Zhipu AI's GLM-5.3-Flash on domestic chips sends stock up 12%
Synopsis
Key Takeaways
Zhipu AI officially unveiled GLM-5.3-Flash — previously code-named Ox Alpha — on Wednesday, 27 August 2026, revealing that the open-weight model ran entirely on a cluster of 100,000 domestically produced chips during a stealth pre-release trial. The disclosure triggered a sharp market response, with Zhipu AI shares closing more than 12 per cent higher at HK$1,160 in Hong Kong on Thursday.
Record-breaking pre-launch traffic
Before its formal debut, Ox Alpha was quietly deployed on AI model marketplace OpenRouter and agent platform OpenCode, where it processed a combined 62 trillion tokens, according to the company. On OpenRouter alone, the model handled more than 11 trillion tokens in its first three days — the largest launch by token volume in the platform's history.
By Thursday, OpenRouter data showed GLM-5.3-Flash ranked first among coding-focused models on the platform, accounting for 10.3 trillion tokens, or nearly 31 per cent of its total weekly volume. That positions it ahead of rival coding assistants from Anthropic, Alibaba's Qwen series, and Moonshot AI's Kimi K lineup on the same leaderboard.
Why it matters: a stress-test for Chinese silicon
The deployment is being closely watched as one of the most demanding real-world trials of China's homegrown semiconductor stack to date. Running inference at this scale — across a 100,000-chip cluster — on domestic hardware, without relying on Nvidia GPUs, directly tests the readiness of processors from suppliers such as Huawei Technologies, Moore Threads, and Cambricon Technologies.
Beijing has made reducing dependence on advanced US processors a strategic priority, particularly as Washington's export controls have progressively restricted China's access to Nvidia's highest-performance chips. A successful large-scale inference run on domestic silicon would mark a meaningful proof-of-concept for that ambition.
Competitive backdrop
The GLM-5.3-Flash launch arrives as the global race to dominate open-weight AI models intensifies. Benchmarking platform Artificial Analysis and marketplace data from OpenRouter increasingly show Chinese models competing directly with Western counterparts on cost-per-token and throughput metrics. Zhipu's consumer and developer platform Z.ai is expected to integrate the new model imminently, expanding its reach beyond API users.
What's next
Investors and industry observers will now watch whether Zhipu AI can sustain the token-processing volumes seen during the stealth trial at commercial pricing, and whether the domestic chip cluster can scale further without performance degradation. The broader question — how quickly China's semiconductor ecosystem can close the gap with Nvidia's H100 and B200 architectures for frontier-model inference — remains the defining variable for the country's AI ambitions heading into 2027.