Why DeepSeek V4 Hasn’t Fully Cut Ties with Nvidia

In reality, the strategic calculus behind open-source AI is too complex to allow simple side-picking.

7 mins read
DeepSeek

When DeepSeek released a preview of V4 on April 24, 2026, one detail consumed the technology world: the model had been natively adapted to run on Huawei’s Ascend 950-series inference chips.

For many observers, especially in Washington, this was the long-awaited signal. The narrative wrote itself: China’s AI champion had finally achieved full-stack independence. The decoupling was complete. Nvidia’s grip on the AI world was loosening.

However, it wasn’t.

A quieter, more telling detail sat buried in the technical documentation: DeepSeek V4 had been validated on both Nvidia GPU and Huawei Ascend NPU platforms. The company, far from celebrating a clean break, had gone out of its way to maintain compatibility with the very ecosystem it was supposedly abandoning. More strikingly, DeepSeek had reportedly refused to give Nvidia and AMD early access to optimize for V4, the standard industry courtesy, yet still ensured the model ran smoothly on CUDA at launch. Indeed, Nvidia published a technical blog on the same day demonstrating V4-Pro running on its Blackwell GB200 and B300 hardware at over 150 tokens per second per user, suggesting the model’s CUDA compatibility required no special intervention.

The question this raises is far more interesting than the simplistic “decoupling” narrative that dominated headlines. Why, with a Chinese alternative approaching viability, does DeepSeek still refuse to cut the cord?

The answer reveals something profound, not just about DeepSeek, but about the strategic wisdom that separates this Chinese AI company’s genuine ambition from simply picking a side in geopolitics.

A Tale of Two Silicon Strategies

Before addressing the “why,” we need clarity on the “what.” The relationship between DeepSeek V4 and its silicon choices is not a binary one. It is a sophisticated, dual-track strategy that serves different goals at different stages of the AI lifecycle.

For training, the computationally brutal process of creating the model itself, the picture remains murky.

DeepSeek’s technical report did not specify which GPUs V4 was trained on. A senior Trump administration official alleged the model was trained on Nvidia’s Blackwell chips smuggled into China. DeepSeek denied this, stating it used H800 GPUs and Huawei Ascend 910C chips. Nvidia called the smuggling claims “far-fetched.”

Meanwhile, Liu Zhiyuan, a computer science professor at Tsinghua University, told MIT Technology Review that DeepSeek appears to have adapted only part of V4’s training process for Chinese chips, and the model may still have been trained mainly on Nvidia hardware. Multiple anonymous sources told the same outlet that Chinese chips remain better suited for inference than training. The truth likely involves a division of labor: the main pre-training run, requiring the greatest stability and scale, may have relied on Nvidia’s mature infrastructure, while other stages incorporated Huawei hardware.

It is engineering pragmatism. When you are training a model with 1.6 trillion total parameters, of which 49 billion are active per token in its mixture-of-experts architecture, stability is not a luxury; it is the difference between a successful run and millions of dollars in wasted compute.

For inference, the process of actually serving the model to users, the picture is dramatically different.

DeepSeek has committed to Huawei’s Ascend 950-series chips for its own inference infrastructure. The company gave Huawei early access to optimize for V4 over several weeks, a privilege pointedly denied to American chipmakers. Huawei said its entire Ascend SuperNode product line was “fully adapted” to V4 for model inference.

However, this transition is not yet complete. DeepSeek acknowledged that V4-Pro would face throughput limitations until the second half of 2026, when Huawei’s Ascend 950PR supernodes are expected to ship at scale. The 950PR chips themselves, while officially launched, are still ramping up production. Huawei claims the Atlas 350 accelerator card powered by the 950PR delivers approximately 2.87 times the FP4 compute performance of Nvidia’s H20, the most powerful chip Nvidia can legally sell in China.

On pricing, DeepSeek V4 is dramatically cheaper than Western alternatives, though the exact ratio depends on the comparison. V4-Pro costs $1.74 per million input tokens and $3.48 per million output tokens at list price, roughly one-sixth to one-seventh the cost of GPT-5.5 or Claude Opus 4.7. With cached input, the gap widens further: V4-Pro costs approximately one-tenth as much as GPT-5.5. The smaller V4-Flash variant is even more aggressive at $0.14 per million input tokens and $0.28 per million output tokens, undercutting virtually every Western model.

So we have a model whose training hardware remains disputed, whose inference is being transitioned to Chinese hardware that has not yet shipped at full scale, and whose pricing aggressively undercuts Western competitors. A hybrid creature. A strategic chimera. The question remains: if Huawei chips are the intended future for inference, the part that actually touches users and generates revenue, why not go all the way and declare full independence?

Here, we arrive at the central insight that most geopolitical analysis misses entirely.

The Open-Source Reality: CUDA Is the Lingua Franca

DeepSeek is an open-source model. The entire value proposition of open-source AI rests on this distributed, global community of contributors and adopters. DeepSeek models have been downloaded more than 75 million times on Hugging Face since January 2025.

And what do those developers overwhelmingly use? Nvidia GPUs running CUDA.

Huawei’s competing CANN software stack has made progress, its new CANN Next release introduces CUDA-compatible programming abstractions that lower migration barriers, but CUDA’s ecosystem still dwarfs it by an order of magnitude. The global developer ecosystem for AI, the researchers publishing papers, the startups building products, the enterprises experimenting with fine-tuning, has been built on CUDA for over a decade.

If DeepSeek were to release V4 in a form that ran only on Huawei’s Ascend NPUs, the model would become a “closed-source project.” The weights might still be downloadable, but practically closed to the vast majority of the world’s AI developers who do not own, and cannot easily acquire, Huawei hardware to run or modify the model. The open-source community would shrug and move on. The model would become a curiosity, a political statement rather than a living, evolving technical artifact.

This explains DeepSeek’s careful balancing act. By ensuring V4 runs on both CUDA and Ascend, the company is speaking two languages simultaneously. To the Chinese ecosystem, they have a viable, high-performance, domestically-controlled option. To the global developer, it says: nothing changes for you. Download, run, modify, build. This dual fluency is the fundamental requirement for remaining an open-source project of global relevance while also serving strategic domestic interests.

The Strategic Calculus: Optionality as Power

By maintaining genuine dual-platform capability, DeepSeek is acquiring something more valuable than any single hardware relationship: optionality.

In strategy, optionality means preserving the freedom to act in multiple ways depending on how circumstances evolve. A company locked into a single supplier, whether Nvidia or Huawei, has no optionality. It is a price-taker, a supplicant, a hostage to another party’s roadmap and fortunes. DeepSeek, by demonstrating it can operate at frontier levels on both major AI hardware ecosystems simultaneously, has made itself valuable to both while being dependent on neither.

Consider the implications. If U.S. export controls tighten further, DeepSeek can shift more workload to Huawei without missing a beat, the inference stack is already being prepared, and portions of the training pipeline have been validated on Ascend hardware. If Huawei’s next-generation chips face yield issues or delays, a real possibility, given that even the 950PR has not yet shipped at scale, DeepSeek can maintain momentum on Nvidia hardware without panicking. If a third ecosystem emerges, perhaps from a European chipmaker seeking relevance, DeepSeek’s demonstrated ability to port efficiently makes it a natural early adopter.

This is not the behavior of a company trying to win favor with Beijing or Washington, but the behavior of a company pursuing strategic autonomy, the freedom to chart its own course regardless of how the winds of geopolitics shift. This very autonomy, this refusal to be anyone’s instrument, may be the most genuinely patriotic posture possible. A DeepSeek that becomes a mere appendage of Huawei’s chip strategy would be a weaker DeepSeek, and ultimately a weaker asset for China’s technological ambitions.

What This Tells Us About DeepSeek’s Real Playbook

The DeepSeek V4 silicon strategy is a Rorschach test. Those who want to see decoupling will see the Huawei adaptation and declare victory. Those who want to see continued dependence will point to the unresolved questions about training hardware and shake their heads. Both miss the point entirely.

What DeepSeek is actually demonstrating is a new model of technological competition, one that is neither pure decoupling nor naive globalism, but something more sophisticated. It is the model of the multi-ecosystem platform. The insight is that, in an era of technological fragmentation, power accrues not to those who pick a side most enthusiastically, but to those who can operate across all sides, translating between them, setting standards that work everywhere, and refusing to be captured by any single infrastructure provider.

This is, in a sense, the ultimate realization of the open-source philosophy. Open-source has always been about freedom from vendor lock-in. DeepSeek is applying that principle at the silicon level. The model does not belong to Nvidia. It does not belong to Huawei. It belongs to the developers who run it, on whatever hardware they choose, for whatever purposes they imagine.

It is important to note, however, that V4 remains a preview release. DeepSeek has not provided a timeline for finalization, and the model’s performance characteristics may evolve before it reaches production-stable status. The Huawei inference story, in particular, is still unfolding, dependent on the 950PR achieving volume production and proving its reliability under sustained, large-scale workloads.

The American analysts who frame every development in the AI industry as a chapter in a zero-sum struggle between the U.S. and China will struggle to grasp this. Their framework demands that every actor choose a side. But DeepSeek V4’s continued embrace of CUDA, even amid enthusiastic adoption of Huawei hardware, is a quiet rebuke to that binary thinking. It is an assertion that the most powerful position in a fragmented technological landscape is not the most loyal ally, but the indispensable broker, the entity that moves fluidly between ecosystems, setting standards that transcend hardware, and serving a global community that refuses to be divided by geopolitics.

The model that runs everywhere will outlast the model that runs only somewhere. DeepSeek seems to understand this.

Source: China Academy

SLG Syndication

SLG Syndication is committed to aggregating excerpts from news published by international news agencies and key insights on contemporary issues published by think tanks. Our aim is to facilitate the expansion of its reach while giving due credit to the original source.

Leave a Reply

Your email address will not be published.

Latest from Blog