China’s race to develop next-generation artificial intelligence is entering a new phase as a shortage of high-quality training data emerges as a potentially critical constraint on its technological ambitions.
While US restrictions on advanced computing chips have dominated the global debate over China’s AI capabilities, Chinese experts are increasingly warning that access to reliable, high-quality data could become an equally significant bottleneck, one that cannot easily be overcome through hardware alternatives.
The challenge is global. US-based research institute Epoch AI estimates that the worldwide supply of high-quality, publicly available human-generated text could be fully exhausted within the next six years. OpenAI co-founder Andrej Karpathy has separately warned of a looming “data wall” by the end of the decade, potentially limiting improvements in model capabilities unless developers can obtain fresh and reliable information.
Leading US AI companies are already searching for new sources of human knowledge. Under an internal initiative code-named Project Panama, Amazon.com-backed Anthropic spent tens of millions of dollars acquiring millions of physical books, severing their bindings and scanning every page before discarding the originals, according to court filings reported by The Washington Post in January. The disclosure prompted criticism from authors, archivists and preservationists, although an Anthropic representative told fact-checking site Snopes earlier this month that “None of our data-acquisition programmes buy and destroy ‘rare’ or ‘antiquarian’ books.”
For China, the pressure is particularly acute because relatively little of the world’s online content is available in Chinese. W3Techs, an internet tracker, says Chinese accounts for 1.3 per cent of global web content as of this month, compared with nearly half for English, 6 per cent for Spanish, 5.9 per cent for German and 5 per cent for Japanese.
Beijing is responding by treating data as a strategic resource. In June, the National Data Administration unveiled a nationwide plan to increase the supply, circulation and commercialisation of high-quality AI training data. By 2028, China aims to establish an extensive ecosystem of validated datasets covering manufacturing, energy, healthcare, finance and agriculture, as well as emerging areas including embodied AI, autonomous driving and low-altitude aviation.
“Competition in the AI era is not only about models and computing power, but also about high-quality data supply systems,” Yu Xiaohui, president of the state-affiliated China Academy of Information and Communications Technology, said in an article published in late July on the National Data Administration’s website.
Chinese scholars and industry experts are consequently calling for a coordinated effort to expand Chinese-language corpora, including large, curated collections of structured text used to train and evaluate AI models. Rather than depending solely on web scraping, they argue that China should systematically digitise extensive offline resources.
Sun Maosong, a computer scientist at Tsinghua University, said in the state-backed Guangming Daily last month that corpora had become a core element in the development and improvement of large language models. He urged authorities to scan historical archives, local gazetteers, ancient manuscripts and scientific literature, while also expanding datasets to include dictionaries, audiovisual material and regional dialects.
China is also promoting simulation and synthetic data generation. A 2024 white paper by AliResearch, a think tank affiliated with Alibaba Group Holding, described algorithms, compute and data as the three pillars of AI model development and said “higher quality and more abundant data” were central to the success of generative AI models.
Yet efforts to digitise offline material are already encountering resistance from publishers and copyright holders. Huaxia Publishing House recently included a warning in a new translation of Huangting Jing and Yinfu Jing, stating: “It is prohibited to use the content of this book for artificial intelligence training. Violators will be held legally responsible.”
Huaxia said the clause had been added at the request of licensing partners for imported titles, while acknowledging that detecting AI infringement and enforcing those rights could be difficult, according to Jiemian News.

