Zhipu AI's GLM-5.3-Flash on domestic chips sends stock up 12%

Share:
Audio Loading voice…
Zhipu AI's GLM-5.3-Flash on domestic chips sends stock up 12%

Synopsis

Zhipu AI's GLM-5.3-Flash — formerly Ox Alpha — processed 62 trillion tokens on a 100,000-unit domestic chip cluster before launch, topped OpenRouter's coding charts with 31% of weekly volume, and sent the company's Hong Kong shares up 12% — the most rigorous public stress-test yet of China's homegrown AI silicon.

Key Takeaways

Zhipu AI launched GLM-5.3-Flash (code-named Ox Alpha ) on 27 August 2026 , powered entirely by a 100,000-unit domestic chip cluster.
The model processed 62 trillion tokens across OpenRouter and OpenCode during its stealth pre-release trial, according to the company.
On OpenRouter , GLM-5.3-Flash set the platform's biggest-ever launch record with more than 11 trillion tokens in its first three days.
By Thursday , the model held the #1 coding model ranking on OpenRouter , capturing nearly 31 per cent of the platform's weekly token volume ( 10.3 trillion tokens ).
Zhipu AI shares closed 12 per cent higher at HK$1,160 in Hong Kong on Thursday, 28 August 2026 .
The deployment represents one of the largest-scale inference workloads run on Chinese domestic chips , amid ongoing US export controls on Nvidia GPUs.

Zhipu AI officially unveiled GLM-5.3-Flash — previously code-named Ox Alpha — on Wednesday, 27 August 2026, revealing that the open-weight model ran entirely on a cluster of 100,000 domestically produced chips during a stealth pre-release trial. The disclosure triggered a sharp market response, with Zhipu AI shares closing more than 12 per cent higher at HK$1,160 in Hong Kong on Thursday.

Record-breaking pre-launch traffic

Before its formal debut, Ox Alpha was quietly deployed on AI model marketplace OpenRouter and agent platform OpenCode, where it processed a combined 62 trillion tokens, according to the company. On OpenRouter alone, the model handled more than 11 trillion tokens in its first three days — the largest launch by token volume in the platform's history.

By Thursday, OpenRouter data showed GLM-5.3-Flash ranked first among coding-focused models on the platform, accounting for 10.3 trillion tokens, or nearly 31 per cent of its total weekly volume. That positions it ahead of rival coding assistants from Anthropic, Alibaba's Qwen series, and Moonshot AI's Kimi K lineup on the same leaderboard.

Why it matters: a stress-test for Chinese silicon

The deployment is being closely watched as one of the most demanding real-world trials of China's homegrown semiconductor stack to date. Running inference at this scale — across a 100,000-chip cluster — on domestic hardware, without relying on Nvidia GPUs, directly tests the readiness of processors from suppliers such as Huawei Technologies, Moore Threads, and Cambricon Technologies.

Beijing has made reducing dependence on advanced US processors a strategic priority, particularly as Washington's export controls have progressively restricted China's access to Nvidia's highest-performance chips. A successful large-scale inference run on domestic silicon would mark a meaningful proof-of-concept for that ambition.

Competitive backdrop

The GLM-5.3-Flash launch arrives as the global race to dominate open-weight AI models intensifies. Benchmarking platform Artificial Analysis and marketplace data from OpenRouter increasingly show Chinese models competing directly with Western counterparts on cost-per-token and throughput metrics. Zhipu's consumer and developer platform Z.ai is expected to integrate the new model imminently, expanding its reach beyond API users.

What's next

Investors and industry observers will now watch whether Zhipu AI can sustain the token-processing volumes seen during the stealth trial at commercial pricing, and whether the domestic chip cluster can scale further without performance degradation. The broader question — how quickly China's semiconductor ecosystem can close the gap with Nvidia's H100 and B200 architectures for frontier-model inference — remains the defining variable for the country's AI ambitions heading into 2027.

Point of View

000 domestic chips, at scale, on a global platform, is precisely the kind of data point Beijing needs to argue that its semiconductor self-sufficiency push is production-ready, not merely aspirational. Mainstream coverage focuses on the OpenRouter rankings, but the more consequential signal is whether Huawei, Moore Threads, or Cambricon silicon can sustain this throughput without the thermal and yield constraints that have historically hobbled Chinese GPU alternatives. Zhipu's 12% single-day stock move also reflects investor repricing of the export-control risk premium — if domestic chips can handle frontier inference, the ceiling on Chinese AI companies rises materially. The real test comes when pricing data emerges: cost-per-token parity with Nvidia-based inference would be the moment the chip-war calculus fundamentally shifts.
NationPress
27 Aug 2026

Frequently Asked Questions

What is Zhipu AI's GLM-5.3-Flash model?
GLM-5.3-Flash is an open-weight AI model released by Zhipu AI on 27 August 2026, previously code-named Ox Alpha during a stealth pre-release trial. It ran on a cluster of 100,000 domestically produced chips and processed 62 trillion tokens before its official launch, setting a record on the OpenRouter platform.
Why did Zhipu AI shares rise 12%?
Zhipu AI shares closed more than 12 per cent higher at HK$1,160 in Hong Kong on Thursday, 28 August 2026, following the formal reveal of GLM-5.3-Flash. Investors responded to the model's record-breaking pre-launch traffic and the demonstration that large-scale AI inference could run on Chinese domestic chips without Nvidia hardware.
How does GLM-5.3-Flash rank on OpenRouter?
As of Thursday, 28 August 2026, GLM-5.3-Flash ranked first among coding models on OpenRouter, accounting for 10.3 trillion tokens — nearly 31 per cent of the platform's total weekly volume. Its three-day pre-release total of more than 11 trillion tokens was the largest launch in OpenRouter's history.
Which domestic chips did Zhipu AI use for GLM-5.3-Flash?
Zhipu AI said the model ran on a cluster of 100,000 domestically produced chips, though the company did not specify the exact chip supplier in its announcement. China's primary domestic GPU and AI accelerator makers include Huawei Technologies, Moore Threads, and Cambricon Technologies.
How does GLM-5.3-Flash fit into the US-China chip competition?
The deployment is a direct response to US export controls that restrict China's access to Nvidia's advanced processors. Running frontier-model inference at scale on domestic silicon is a key benchmark for Beijing's semiconductor self-sufficiency strategy, and a successful rollout like this narrows the perceived capability gap between Chinese and US-supplied AI hardware.
Nation Press
The Trail

Connected Dots

Tracing the thread behind this story — newest first.

8 Dots
  1. Latest Yesterday
  2. 1 week ago
  3. 3 weeks ago
  4. 1 month ago
  5. 1 month ago
  6. 1 month ago
  7. 2 months ago
  8. 2 months ago
Google Prefer NP
On Google