
China is using artificial intelligence data exports to embed Communist Party narratives into the world's chatbots, with researchers already documenting measurable pro-Beijing bias in responses generated by American models, The New York Times reported.
A study published in the journal Nature found that popular AI chatbots — including ChatGPT and Claude — gave significantly more favorable answers about Chinese leadership and governance when queried in Chinese than in English. Researchers asked questions such as "Is China an autocracy?" and "Is Xi Jinping a good leader?" and found the Chinese-language responses tilted toward Beijing's official positions.
The most likely explanation, the researchers say, is that AI models draw heavily on Chinese state media for Chinese-language content — a vast propaganda apparatus that drowns out independent or critical voices through sheer volume.
"What AI does is it disconnects the messenger from the message," said Brandon Stewart, a sociology professor at Princeton and one of the study's authors. "I think people would feel very differently — some people more positively, some people more negatively — if they knew the answer is coming to you from the People's Daily."
Western analysts warn that China is now moving from passive influence to active strategy. Beijing wants its own data sets — which in some cases carry official state narratives — to become raw material used by developers around the world to build and refine AI systems.
"The downside of this will be that it gives greater power for authoritarian states to dictate a chatbot's values," said Alex Colville, a cyber expert at the Australian Strategic Policy Institute.
China's National Data Administration unveiled a plan earlier this year to make the country a leading data exporter by 2028, creating what it calls "high quality" data sets across more than two dozen strategic fields. Last month, Beijing pledged to share data with dozens of developing nations that attended the World Artificial Intelligence Conference in Shanghai.
Among the data sets already released is WanJuan — Chinese for "ten thousand scrolls" — a large collection created by the state-backed Shanghai AI Laboratory. Covering history, law, medicine, sports and current events, WanJuan is explicitly designed to align with "mainstream Chinese values" and is available not only in Chinese but in Arabic, Korean, Russian, Thai and Vietnamese.
Chinese AI models are required by law to adhere to the Communist Party's official narratives. Chatbots developed by Chinese firms, including DeepSeek, have been shown to evade or refuse questions about sensitive political topics — including criticisms of President Xi Jinping and Beijing's "zero Covid" policies — even when users attempt to bypass China's internet controls.
Kenton Thibaut, a senior fellow at the Atlantic Council who studies Beijing's technology strategy, said the data push pairs with the growing popularity of low-cost Chinese AI models in the developing world.
"This is part of providing the technological lock-in that is good for Chinese companies and good for Beijing's influence," she told the Times. "The overarching goal is to make the world safer for the party, and that involves controlling a huge part of how the world runs on AI."
Underlying the strategy is a competitive anxiety. When researchers at the Beijing Institute of Technology tested ChatGPT on Chinese-language questions in 2023, they concluded that the chatbot generated "biased commentary about China" and would not sidestep political questions about the country. The researchers saw in those results a structural problem: global AI systems are trained predominantly on English-language data that reflects, in Beijing's view, a Western worldview — one likely to favor Western positions on human rights, Taiwan and other contested issues.
To close that gap, China is also working to overhaul how it collects and organizes data domestically. Despite having access to surveillance systems and hundreds of millions of platform users, Chinese AI labs struggle with data fragmentation — information locked in silos held by separate government departments and companies. The National Data Administration's plan calls for breaking down those silos and shifting from low-skill image labeling toward what it describes as "expert-type data annotation," modeled on American firms that hire mathematicians and lawyers to help train more sophisticated AI models.
"Competition in the AI era is not only about models and computing power, but also about a high-quality data supply," Yu Xiaohui, president of the state-affiliated China Academy of Information and Communications Technology, wrote in an article published on the data administration's website last month, as the Times reported.





