As China seeks greater influence over the development of artificial intelligence, Beijing is confronting a strategic problem: the systems shaping the technology’s future are trained predominantly on English-language data and can reflect what Chinese authorities regard as a Western way of thinking. The effort to become a leading supplier of the data used to train A.I. is therefore becoming central to China’s ambitions in the field.
When ChatGPT was still a relatively new technology, researchers in Beijing put the chatbot through a series of tests designed to assess how well it handled Chinese-language questions. The results, published in 2023 by researchers at the Beijing Institute of Technology, offered an early indication of the tensions surrounding the use of artificial intelligence in China.
Among the examples cited by the researchers was a basic factual error involving Yao Ming, the former N.B.A. star. ChatGPT described him as the first Chinese woman to play professional basketball in the United States. The chatbot also confused two of China’s classic works of literature, “Journey to the West” and “Dream of the Red Chamber”, which were written two centuries apart.
The researchers reported problems beyond factual mistakes. They wrote that ChatGPT generated a large amount of “biased commentary about China” and “would not evade or refuse to answer political questions about China”.
ChatGPT has since undergone many updates, meaning it is unclear whether the same questions would produce the same results today. Yet the examples highlighted a broader concern that extends beyond the performance of any single chatbot. For China, the issue is not simply whether artificial intelligence can correctly understand Chinese language and culture. It is also about whose information, perspectives and assumptions are embedded in the systems increasingly shaping how people interact with knowledge.
The underlying imbalance, according to analysts, is rooted partly in the data on which artificial intelligence systems are trained. Much of that material is in English, giving the technology a foundation that China sees as reflecting a Western way of thinking. In a field where the quality and breadth of training data are fundamental to the development of increasingly powerful systems, the composition of that data has therefore become a strategic concern.
The implications are particularly sensitive for the Chinese Communist Party. Analysts say the dominance of Western-oriented data creates a vulnerability because Western views are likely to prevail on issues such as human rights and the status of Taiwan, the self-governed island claimed by Beijing.
That concern gives China’s A.I. ambitions another dimension. The contest is not only about producing increasingly powerful chatbots or developing the technology capable of running them. It is also about the vast quantities of information that make those systems possible: the text, images and videos used to train artificial intelligence.
Beijing therefore wants China to become a leading supplier of data for the development of A.I. systems around the world. Such an ambition would address what Chinese authorities and analysts see as a fundamental imbalance in the technological ecosystem while giving China a greater role in determining the information available to future artificial intelligence systems.
The significance of that objective can be seen against the backdrop of China’s broader effort to establish itself as an artificial intelligence power. At the World Artificial Intelligence Conference in Shanghai last month, a Google booth stood among numerous displays, with Chinese text visible alongside the company’s logo. Other exhibitors included KIMI LAB, reflecting the increasingly international and competitive environment surrounding the technology in China.
The presence of international and Chinese technology interests at such an event illustrates the extent to which artificial intelligence has become an arena in which questions of technology and information intersect. For China, the challenge is not merely to participate in the development of systems created elsewhere, but to increase its influence over the material from which those systems learn.
The early experience with ChatGPT demonstrated why that objective matters to Beijing. When a chatbot misidentified Yao Ming and confused two foundational works of Chinese literature, the errors were visible and straightforward. More consequential, from China’s perspective, was the researchers’ assessment that the system generated biased commentary about China and did not refuse political questions concerning the country.
Those findings were published in 2023, when ChatGPT was still new, and they cannot establish how current systems would respond to the same questions. But they provide a window into the concerns surrounding artificial intelligence in China: language, culture, political questions and the composition of training data are closely connected.
The debate is consequently moving beyond the familiar question of which country can build the most capable artificial intelligence system. It is also becoming a question of who supplies the information from which those systems learn and whose perspectives are represented within that information.
For Beijing, becoming a major supplier of data could help address a technological disadvantage created by the predominance of English-language material. It could also give China a greater voice in the development of A.I. chatbots and other systems whose influence extends across national boundaries.
The stakes are therefore larger than the accuracy of an individual chatbot. Artificial intelligence systems increasingly depend on enormous collections of text, images and videos, and the sources represented in those collections can influence what the systems know, how they interpret questions and what perspectives they reproduce.
China’s effort to become a leading supplier of such data reflects its recognition of that reality. As artificial intelligence develops, control over the information that trains the technology may prove nearly as important to Beijing as the technology itself. For a country concerned that Western views dominate the systems shaping the future, increasing its role in supplying the data offers a way to address what it sees as a strategic vulnerability while seeking a greater voice in the direction of artificial intelligence.

