Chinese artificial intelligence start-up DeepSeek has drawn fresh attention from the global tech community after its founder, Liang Wenfeng, co-authored a technical paper outlining a novel approach to training large language models that could significantly reduce dependence on scarce and expensive computing resources. The research, completed with scholars from Peking University and released on January 12, proposes a method that allows aggressive scaling of AI model parameters while bypassing traditional GPU memory constraints, a long-standing bottleneck in AI development.
The move highlights DeepSeek’s continued emphasis on efficiency-driven innovation as it competes with far better-funded US rivals. According to a report published the following day, the development has fuelled market speculation that the company may soon unveil a major new AI model, possibly before the Lunar New Year. The paper is already being closely examined by industry experts in both China and the United States, eager to assess the implications for large-scale model training.
Titled “Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models,” the paper introduces a technique known as Engram, which applies a form of conditional memory to large models. The authors argue that current AI systems waste enormous computing power by repeatedly recalculating basic information instead of efficiently storing and retrieving it. This inefficiency consumes GPU memory and limits the depth available for more advanced reasoning tasks.
By separating, or “decoupling,” computation from storage, Engram enables models to retrieve foundational knowledge through rapid lookups rather than repeated calculations. This approach significantly reduces pressure on high-bandwidth memory, an area where China continues to lag behind the United States and other leading semiconductor producers. Analysts note that while Chinese firms such as ChangXin Memory Technologies have made progress, they remain several years behind industry leaders including Samsung Electronics, SK Hynix and Micron Technology.
The researchers also say the technique improves a model’s ability to handle long-context inputs, a crucial capability for deploying AI systems as real-world agents rather than simple chatbots. Tests conducted on a 27-billion-parameter model showed performance gains of several percentage points across major industry benchmarks, while preserving more computational capacity for complex reasoning. The authors described conditional memory as a foundational building block for the next generation of sparse AI models, comparing its potential impact to DeepSeek’s earlier work on Mixture-of-Experts architectures.
Industry reaction has been swift. Elie Bakouch, a research engineer at open-source AI platform Hugging Face, praised the paper for validating the approach on real hardware during both training and inference. The research team includes 14 co-authors, among them Zhang Huishuai, an assistant professor at Peking University and a former senior researcher at Microsoft Research Asia, underscoring the academic and industry pedigree behind the work.
DeepSeek rose to international prominence early last year with the release of its DeepSeek-R1 model, which was trained using Nvidia H800 GPUs at a reported cost of just US$5.5 million. The model achieved performance comparable to leading US systems despite being developed in a fraction of the time and at a small fraction of the cost, drawing particular attention from American policymakers and technology executives.
The company’s growing influence was highlighted again this month when Microsoft president Brad Smith warned that US AI firms are losing ground to Chinese competitors in markets outside the West. He cited DeepSeek as a prime example of how low-cost, open-source models from China are gaining rapid adoption in regions such as Africa. A Microsoft study found that DeepSeek’s R1 model helped accelerate AI uptake across the Global South, contributing to China overtaking the United States in global market share for open-source AI models.
As the first anniversary of the R1 model approaches, expectations are building that DeepSeek will soon make another major announcement. Industry observers point to reports suggesting the company could release a powerful new V4 model with advanced programming capabilities as early as mid-February. If confirmed, such a launch would further cement DeepSeek’s position as a key challenger in the global AI race, particularly as efficiency and accessibility become increasingly decisive factors in determining market leadership.

