DeepSeek’s AI Model R1 Revealed in First Peer-Reviewed Study

DeepSeek also addressed speculation that R1’s success relied on outputs from rival models like OpenAI’s.

1 min read
DeepSeek founder Liang Wenfeng [File Image]

A peer-reviewed paper published in Nature has shed new light on how Chinese start-up DeepSeek built its headline-grabbing artificial intelligence model R1 for a fraction of the cost of its rivals.

R1, which rattled U.S. stock markets when released in January, was designed to excel at reasoning tasks such as mathematics and coding. Unlike many Western large language models (LLMs), R1 is available as an “open weight” system, making it free to download and widely used. According to Hugging Face, it has already been downloaded more than 10.9 million times — making it the platform’s most popular open model.

The Nature paper revealed for the first time that DeepSeek spent only around US$294,000 training R1, on top of roughly $6 million invested in building its base LLM. This contrasts sharply with the tens of millions of dollars typically needed to train competing models. The company trained R1 largely on Nvidia’s H800 chips, which were restricted from sale to China under U.S. export controls.

What sets R1 apart is its training approach. Rather than relying on human-annotated reasoning examples, DeepSeek applied a form of pure reinforcement learning. The model was rewarded for producing correct answers and developed its own reasoning-like strategies, including verifying solutions independently. To make the process more efficient, R1 even scored its own outputs, using a method called group relative policy optimization.

“This is a very welcome precedent,” said Lewis Tunstall, a machine-learning engineer at Hugging Face who reviewed the paper. “If we don’t have this norm of sharing a large part of this process publicly, it becomes very hard to evaluate whether these systems pose risks or not.”

DeepSeek also addressed speculation that R1’s success relied on outputs from rival models like OpenAI’s. The company stated in exchanges with referees that R1 did not directly train on OpenAI-generated reasoning data, though like other LLMs, its base model drew from the open web — which inevitably included some AI-generated content.

Huan Sun, an AI researcher at Ohio State University, said R1 has been “quite influential” in shaping reinforcement learning approaches for LLMs in 2025. Other labs are now adapting its methods to strengthen reasoning abilities across domains beyond math and coding.

Although R1 was not the most accurate model in scientific tasks such as those benchmarked by ScienceAgentBench, it was praised for offering an exceptional balance of performance and cost. “R1 has kick-started a revolution,” Tunstall said.

Sri Lanka Guardian

The Sri Lanka Guardian is an online web portal founded in August 2007 by a group of concerned Sri Lankan citizens including journalists, activists, academics and retired civil servants. We are independent and non-profit. Email: editor@slguardian.org

Leave a Reply

Your email address will not be published.

Latest from Blog