China's top AI models still run on Nvidia chips despite local push
Synopsis
Key Takeaways
China's most advanced large language models (LLMs) continue to be trained on Nvidia chips, according to sources at major Chinese AI developers, as the prohibitively high cost of migrating to domestic semiconductors stalls Beijing's self-sufficiency drive. The disclosure, surfacing on 10 August 2026, underscores a stubborn dependency that persists even as homegrown hardware makers race to close the gap.
The Software Lock-In Problem
At the heart of the delay is a deep-rooted software ecosystem challenge. Nvidia's Compute Unified Device Architecture (CUDA) platform has long been the de facto standard for AI model development worldwide. Switching away from it is not simply a matter of swapping hardware — it demands a wholesale rewrite of training pipelines, tooling, and optimisation logic.
'Training LLMs on Nvidia chips for now remains the norm among Chinese AI developers,' said a person familiar with the industry. The statement reflects a consensus that has quietly persisted even as geopolitical pressure mounts on Chinese firms to reduce reliance on US technology.
Huawei's CANN vs Nvidia's CUDA
Huawei Technologies' alternative compute platform — Compute Architecture for Neural Networks (CANN), designed for its Ascend chip series — requires developers to rewrite and optimise large amounts of existing code, according to an AI researcher involved in model development. This is not a minor adjustment; it represents a fundamental re-engineering of workflows that teams have spent years building on CUDA.
James Wang, who develops AI models at a research institute affiliated with a Shanghai-based university, put it plainly: 'Our existing training pipelines are reliant on CUDA. CUDA code cannot run directly on Ascend and requires extensive rewriting.' Wang estimated that migrating existing workflows to Huawei's Ascend chips could add at least 50 per cent in time and costs for his team.
Who Is Affected
The dependency spans some of China's most prominent AI players. Developers behind models including DeepSeek, Kimi K3 by Moonshot AI, LongCat, and products from Alibaba Group Holding and Meituan are all navigating the same bottleneck, according to industry sources. For these companies, the engineering overhead of a full chip migration is not merely inconvenient — it is commercially prohibitive in the near term.
The Competitive Backdrop
The situation highlights a structural asymmetry in the global AI chip race. While Huawei's Ascend series has made measurable hardware progress, software maturity — compilers, libraries, debugging tools, and community support — remains years behind CUDA's entrenched ecosystem. Domestic alternatives lack the breadth of third-party optimisations that Nvidia's platform has accumulated over more than a decade.
The US export control regime has restricted China's access to Nvidia's most advanced chips, yet Chinese developers continue to work with available Nvidia hardware rather than pivot to local silicon, signalling that regulatory pressure alone has not been sufficient to force a transition.
What's Next
The pace at which Huawei and other domestic chipmakers can mature their software stacks will be the decisive variable. Until CANN and comparable platforms can absorb CUDA-native workflows with minimal friction, China's frontier AI development will remain tethered to Nvidia infrastructure — a dependency that both complicates Beijing's technology ambitions and exposes leading AI developers to ongoing supply-chain risk.