-
Chinese police researchers developed an AI model to translate and assess Chinese-language dark web content.
-
The team collected over 10,000 records from eight Chinese-language trading sites and forums.
-
The model classified 12,467 samples into eight groups, including illegal services, stolen data, and violent content.

Chinese researchers linked to law enforcement have revealed how AI can help study hidden dark web content. The team worked on a system that could collect, translate, and classify material from Chinese-language dark websites.
The research offers a rare look at how police-linked scientists studied a part of the internet that is hard to access. The dark web uses encrypted connections and anonymous routing systems. These tools make it harder to track users and their online activity.
Criminals also use the dark web to trade stolen data and offer illegal services. Researchers in China and other countries have studied dark web activity for years. However, much of that work has focused on English-language content.
Researchers have published far less work about Chinese-language dark web material. In May, a team from the Criminal Investigation Police University of China published a study about the issue. The team published the paper in the Journal of Chinese Computer Systems. The researchers created a model made for Chinese dark web content.
Researchers Build Tools to Collect Hidden Data
The team first faced a major problem. They needed to enter dark net sites and collect information from them. The researchers found that different sites used different login systems. Traditional web crawlers also struggled to pass Captcha checks.
The sites also changed their anti-scraping methods often. That made automatic data collection more difficult. The researchers built a system that could perform several tasks. These included logging in, creating lists, and downloading images. The system also controlled how often it accessed each site. This helped it avoid detection systems used by the sites. The research team eventually collected over 10,000 data records.
The technological development mirrors similar work in South Korea, where authorities have developed advanced tools to hunt dark web drug dealers.
The data came from eight (8) Chinese-language forums and trading sites, according to the paper. The team then sent the text and images to another model for processing. That model changed images into standard text descriptions.
The images included material such as bank card details. The researchers then manually labelled 1,800 samples. They used these samples as starting data to train the first classification model. The model later used active learning to label the remaining data. Human reviewers then checked the results. After six rounds, the model had classified 12,467 samples.
AI Model Sorts Content by Risk
The model placed the samples into eight different categories. These categories included financial credential trading and the trafficking of identity data. They also covered illegal services, sexually explicit material, and extreme or violent content.
The researchers also divided suspicious material into two levels. The levels were minor violations and serious crimes. The team created a scoring system because much of the collected material carried some level of risk.
Content without sensitive or illegal material received scores from 0 to 0.2. Material with sensitive phrases but no illegal content received scores between 0.2 and 0.4. Illegal content received scores between 0.4 and 1.
The final score depended on what the material contained. The system added 0.05 points for identity theft or offers to provide illegal services. References to minors or criminal knowledge and tools added another 0.1 points. The researchers first created an inventory of sensitive phrases by hand.
They based the list on phrases and words connected to illicit activity on the dark web. The team later updated the word list using predictions from the model. However, the paper did not explain how researchers made those updates. The researchers tested the finished model with an independent group of 200 participants. The test found that the AI tools could reliably classify content found on the dark web.
Researchers Outline Future Plans
The paper did not say whether law enforcement agencies would use the model to monitor Chinese dark web posts. The researchers instead focused on the model’s ability to reduce the amount of human work needed. According to the team, the method can perform risk labelling in a stable way while using fewer human resources.
The researchers also said the system could provide reusable data and methods for studying risks on the Chinese dark web. The team wrote that future research would increase the amount of data used by the model. The researchers also plan to explore other analysis tasks designed for the Chinese dark web environment.
The study shows how researchers are using AI to handle large amounts of hidden online content. The system combines automated data collection, image-to-text processing, machine learning, and human checks. The researchers used several rounds of training and review to improve the model’s ability to sort the material.
The paper’s findings also show the challenges involved in studying Chinese-language dark web activity. The team had to deal with changing login systems, Captcha checks, and anti-scraping methods. It then had to process both text and images before the model could classify the material.
The research team said its method could support future studies of risks linked to Chinese dark web content. However, the paper did not confirm any planned use of the model by law enforcement agencies. The researchers’ future plans remain focused on expanding the data and finding more ways to study the hidden online environment.