Chinese AI firms lean on Nvidia for coding as domestic chips fall short
Synopsis
Key Takeaways
Chinese AI companies are being forced to stretch a limited pool of Nvidia processors to handle high-value inference workloads — particularly coding tasks — because domestic chips cannot yet meet the performance bar that enterprise customers demand, according to industry insiders. The constraint is sharpening as China's AI sector pivots from model training to large-scale deployment, exposing a critical gap in the country's semiconductor self-sufficiency drive.
A market splitting in two
Guan Jiawei, vice-president of inference optimisation start-up Approaching.AI, described a deepening divide in the inference market. “The demand side is now showing a bipolarisation,” he said, noting that demand for high-quality tokens — the basic units of data that models process and generate — far outstripped supply. High-tier tasks required stringent performance metrics that domestic processors could not yet reliably deliver, he added.
Guan was pointed about where the bottleneck sits: “If we rely solely on domestic chips for inference, they can only handle the low-quality tier — the tier with weak demand and weak monetisation. That makes it very hard to find a viable commercial path. That’s why high-quality tokens still depend on Nvidia.” Advanced Chinese models, he said, “place high demands on chips … especially in scenarios like coding, where users are willing to pay a premium.”
Why it matters: inference is the new battleground
Unlike training — which is a one-time, compute-intensive process — inference runs continuously every time a user interacts with a model, making it the dominant driver of ongoing chip demand. Chinese firms have made progress adapting inference workloads to domestic hardware such as Huawei Technologies' Ascend 910B, but complex, latency-sensitive tasks like code generation still require the raw throughput of Nvidia processors, including the H20 — a chip specifically configured for the Chinese market under US export controls.
The stakes are rising fast. China's average daily token calls exceeded 140 trillion in March, up more than 1,000-fold from the beginning of 2024, according to the National Data Administration. That explosive growth is being driven partly by agentic AI — systems that perform real-world tasks autonomously rather than simply answering questions — which is far more compute-hungry than conventional chatbot interactions.
The competitive backdrop
The inference crunch lands at a sensitive moment for China's domestic chip ecosystem. Huawei's Ascend series and other homegrown accelerators have been positioned as substitutes for Nvidia hardware, but the performance gap on precision workloads remains a structural challenge. Optimisation frameworks such as KTransformers are helping squeeze more throughput from available silicon, yet software workarounds have limits when the underlying hardware cannot match the memory bandwidth and floating-point performance of restricted Nvidia models like the H200.
The situation also reflects a broader dynamic: models from players including DeepSeek, Moonshot (maker of Kimi K3), and others have demonstrated impressive benchmark results, but converting those benchmarks into profitable, high-quality commercial inference at scale still requires chips that China cannot freely import.
What’s next
Industry observers will be watching whether successive generations of Huawei Ascend silicon can close the coding-inference gap, and whether US export-control tightening further constrains the Nvidia H20 supply that Chinese AI firms currently depend on. Software-layer optimisation will continue, but the monetisation ceiling for domestic-chip-only inference stacks remains a structural risk for the sector’s commercial ambitions.