Huawei Technologies has unveiled a major advancement in artificial intelligence training, claiming it has surpassed leading industry approaches—such as those by DeepSeek—with a novel architecture optimized for its self-developed Ascend chips. The development marks another step in the Chinese tech giant’s bid to reduce reliance on U.S. technology amid ongoing sanctions.
In a research paper published last week, the team behind Huawei’s Pangu large language model (LLM) introduced Mixture of Grouped Experts (MoGE), a refined version of the widely used Mixture of Experts (MoE) architecture. The paper, co-authored by 78 researchers including 22 core contributors, highlights how MoGE overcomes a key limitation of MoE: inefficient expert activation and workload imbalance across multiple devices.
MoE techniques enable large AI models to scale by selectively activating specialized sub-networks—referred to as “experts”—based on input data. While this leads to efficiency in theory, Huawei researchers noted that MoE often causes computational bottlenecks in real-world training and inference due to uneven expert utilization.
Huawei’s MoGE innovation addresses this by grouping experts during selection, which, according to the paper, “better balances the expert workload” and enhances both training and inference efficiency when deployed across many devices in parallel.
The new architecture was tested on Huawei’s Ascend Neural Processing Units (NPUs)—domestically developed hardware designed to accelerate AI workloads. The results were notable: the Pangu model, powered by MoGE, outperformed competitors including DeepSeek-V3, Alibaba’s Qwen2.5-72B, and Meta’s LLaMA-405B, achieving state-of-the-art results on general English benchmarks and leading performance on all Chinese language benchmarks. It also demonstrated superior efficiency in handling long-context tasks.
Huawei’s LLM, Pangu Ultra, with 135 billion parameters, was specifically optimized for NPU architecture. According to the paper, the model underwent a rigorous training process, including:
- Pre-training on 13.2 trillion tokens
- Long-context extension using 8,192 Ascend chips
- A final post-training optimization stage
The model excels in reasoning and language comprehension tasks, making it a potential contender in enterprise and national-level applications. Researchers also confirmed that the Pangu system will soon be available to Huawei’s commercial customers, underscoring the company’s push to commercialize its homegrown AI ecosystem.
This progress comes as the U.S. continues to restrict Chinese firms from accessing advanced semiconductors such as Nvidia’s AI chips. In that context, Huawei’s advancements in both algorithm design and chip development could offer China a significant technological hedge.
As Chinese AI firms double down on integrated software-hardware development to circumvent foreign chip dependency, Huawei’s MoGE-enhanced Pangu model could represent a pivotal shift—not just in performance, but in strategic self-reliance.

