Chinese AI firms lean on Nvidia for coding as domestic chips fall short

Share:
Audio Loading voice…
Chinese AI firms lean on Nvidia for coding as domestic chips fall short

Synopsis

China's AI boom is hitting a hard ceiling: domestic chips can handle routine inference but fail on high-value coding tasks, forcing firms to hoard scarce Nvidia processors even as daily token calls surpass 140 trillion — a 1,000-fold surge since early 2024.

Key Takeaways

Guan Jiawei , vice-president of Approaching.AI , says Chinese domestic chips can only handle low-quality inference tiers, making a viable commercial path difficult without Nvidia hardware.
China's average daily token calls exceeded 140 trillion in March 2026 , up more than 1,000-fold from the start of 2024 , according to the National Data Administration .
Complex tasks such as coding require the throughput of Nvidia chips; domestic alternatives including Huawei Technologies ' Ascend 910B cannot yet reliably meet those performance metrics.
The shift from model training to large-scale agentic deployment is accelerating compute demand, widening the gap between available supply and high-quality inference needs.
Software optimisation frameworks like KTransformers are being deployed to stretch existing chip capacity, but are not sufficient to close the hardware performance gap on premium workloads.

Chinese AI companies are being forced to stretch a limited pool of Nvidia processors to handle high-value inference workloads — particularly coding tasks — because domestic chips cannot yet meet the performance bar that enterprise customers demand, according to industry insiders. The constraint is sharpening as China's AI sector pivots from model training to large-scale deployment, exposing a critical gap in the country's semiconductor self-sufficiency drive.

A market splitting in two

Guan Jiawei, vice-president of inference optimisation start-up Approaching.AI, described a deepening divide in the inference market. “The demand side is now showing a bipolarisation,” he said, noting that demand for high-quality tokens — the basic units of data that models process and generate — far outstripped supply. High-tier tasks required stringent performance metrics that domestic processors could not yet reliably deliver, he added.

Guan was pointed about where the bottleneck sits: “If we rely solely on domestic chips for inference, they can only handle the low-quality tier — the tier with weak demand and weak monetisation. That makes it very hard to find a viable commercial path. That’s why high-quality tokens still depend on Nvidia.” Advanced Chinese models, he said, “place high demands on chips … especially in scenarios like coding, where users are willing to pay a premium.”

Why it matters: inference is the new battleground

Unlike training — which is a one-time, compute-intensive process — inference runs continuously every time a user interacts with a model, making it the dominant driver of ongoing chip demand. Chinese firms have made progress adapting inference workloads to domestic hardware such as Huawei Technologies' Ascend 910B, but complex, latency-sensitive tasks like code generation still require the raw throughput of Nvidia processors, including the H20 — a chip specifically configured for the Chinese market under US export controls.

The stakes are rising fast. China's average daily token calls exceeded 140 trillion in March, up more than 1,000-fold from the beginning of 2024, according to the National Data Administration. That explosive growth is being driven partly by agentic AI — systems that perform real-world tasks autonomously rather than simply answering questions — which is far more compute-hungry than conventional chatbot interactions.

The competitive backdrop

The inference crunch lands at a sensitive moment for China's domestic chip ecosystem. Huawei's Ascend series and other homegrown accelerators have been positioned as substitutes for Nvidia hardware, but the performance gap on precision workloads remains a structural challenge. Optimisation frameworks such as KTransformers are helping squeeze more throughput from available silicon, yet software workarounds have limits when the underlying hardware cannot match the memory bandwidth and floating-point performance of restricted Nvidia models like the H200.

The situation also reflects a broader dynamic: models from players including DeepSeek, Moonshot (maker of Kimi K3), and others have demonstrated impressive benchmark results, but converting those benchmarks into profitable, high-quality commercial inference at scale still requires chips that China cannot freely import.

What’s next

Industry observers will be watching whether successive generations of Huawei Ascend silicon can close the coding-inference gap, and whether US export-control tightening further constrains the Nvidia H20 supply that Chinese AI firms currently depend on. Software-layer optimisation will continue, but the monetisation ceiling for domestic-chip-only inference stacks remains a structural risk for the sector’s commercial ambitions.

Point of View

Revenue-linked, and impossible to defer. The admission from Approaching.AI's Guan Jiawei that domestic chips can only serve the 'weak demand, weak monetisation' tier is a rare candid acknowledgement that China's AI commercialisation roadmap has a hard dependency on restricted Nvidia silicon. Mainstream coverage tends to focus on benchmark parity between models like DeepSeek or Kimi K3 and Western counterparts, but benchmark performance and inference economics at scale are different problems. The 1,000-fold token-usage surge since early 2024 means this constraint will intensify faster than most supply-side forecasts anticipated, putting pressure on both Huawei's next-generation Ascend roadmap and US policymakers weighing further H20 restrictions.
NationPress
20 Aug 2026

Frequently Asked Questions

Why do Chinese AI companies still need Nvidia chips for inference?
Chinese AI companies need Nvidia chips because domestic processors cannot yet meet the performance requirements for high-value inference tasks like coding, where users pay a premium for low-latency, high-accuracy responses. According to Guan Jiawei of Approaching.AI , relying solely on domestic chips limits firms to a low-quality inference tier with weak demand and weak monetisation.
What is the difference between AI training and inference in terms of chip requirements?
AI training is a one-time, compute-intensive process that builds a model, while inference is the ongoing process of the trained model responding to user queries — and it runs at massive scale continuously. Inference can be partially adapted to domestic hardware for simpler tasks, but complex workloads like coding still require the memory bandwidth and throughput of high-end Nvidia chips.
How fast is China's AI token usage growing?
China's average daily token calls exceeded 140 trillion in March 2026 , a more than 1,000-fold increase from the beginning of 2024 , according to the National Data Administration . The surge is being driven by the rise of agentic AI systems that autonomously perform real-world tasks, which are far more compute-intensive than simple question-and-answer interactions.
Can Huawei's Ascend chips replace Nvidia for AI inference in China?
Huawei Technologies ' Ascend 910B can handle lower-complexity inference workloads, but industry insiders say it cannot yet reliably deliver the stringent performance metrics required for high-tier tasks like coding. The gap between domestic alternatives and restricted Nvidia models such as the H200 and H20 remains a structural challenge for China's AI commercialisation.
What is KTransformers and how does it help Chinese AI firms?
KTransformers is a software optimisation framework being used by Chinese AI firms to extract more inference throughput from available chip capacity. While it helps stretch scarce hardware resources, software-layer optimisation has limits and cannot fully compensate for the underlying performance gap in domestic silicon on premium workloads.
Nation Press
The Trail

Connected Dots

Tracing the thread behind this story — newest first.

8 Dots
  1. Latest 6 days ago
  2. 3 weeks ago
  3. 1 month ago
  4. 1 month ago
  5. 1 month ago
  6. 2 months ago
  7. 2 months ago
  8. 3 months ago
Google Prefer NP
On Google