Ox Alpha Unmasked: GLM-5.3-Flash Lifts Zhipu Shares
SCMP reports the model ran on 100,000 Chinese-made chips and moved 62 trillion tokens before launch; the shares closed at HK$1,160.
Illustration: a vast data center hall where rows of server racks recede into cool, hazy light.
Ox Alpha is GLM-5.3-Flash — Zhipu AI has released the viral stealth-test model with open weights, the South China Morning Post reports the trial ran entirely on 100,000 domestically produced chips, and the Hong Kong-listed stock closed more than 12 percent higher at HK$1,160 on Thursday.
At a glance
- Architecture per SiliconANGLE: mixture of experts, 320 billion total parameters with 18 billion active.
- Context: 1 million tokens of input across text, images and video; output up to 131,072 tokens.
- Pricing per Product Hunt: $0.15 per million input tokens, $0.50 per million output tokens, 50% off for two weeks.
- Benchmarks: top score on GDPval-AA v2, second on AutomationBench — measured against Claude Opus 4.8 and GPT-5.6 Terra.
- Weights on Hugging Face under zai-org/GLM-5.3-Flash; OpenRouter runs a free hosted version.
Zhipu AI has published the model that circulated during testing under the code name Ox Alpha. It ships as GLM-5.3-Flash, and the South China Morning Post reports the stealth trial ran entirely on a cluster of 100,000 domestically produced chips. The company's Hong Kong-listed stock closed more than 12 percent higher at HK$1,160 on Thursday, August 27, 2026.
The traffic that gave it away
SCMP puts the volume at 62 trillion tokens processed before Wednesday's formal release. OpenRouter counted more than 11 trillion tokens in the first three days, which the platform calls its largest launch so far. The model also ranks first among coding systems there.
The same report cites a weekly figure of 10.3 trillion tokens, or nearly 31 percent of the platform's total weekly volume. Those two numbers cover different windows, and the report does not reconcile them. We keep them separate rather than adding them together.
Inside the model
SiliconANGLE describes a mixture-of-experts design with 320 billion total parameters and 18 billion active per request. Input runs to 1 million tokens across text, images and video, while output goes up to 131,072 tokens. Training used 30 trillion tokens with an optimization method the report calls mHC.
Sparse attention and linear attention are meant to cut hardware overhead, and a leaner algorithm stands in for the softmax function. SiliconANGLE puts cost efficiency at ten times that of the previous model.
Scores, price and where to get it
GLM-5.3-Flash takes the highest score on GDPval-AA v2, an evaluation of knowledge work, and second place on AutomationBench, which measures task completion in cloud applications. The comparison set was Claude Opus 4.8, GPT-5.6 Terra and Gemini 3.7 Flash. Weights sit on Hugging Face under zai-org/GLM-5.3-Flash, and OpenRouter operates a free hosted version.
Product Hunt lists API pricing at $0.15 per million input tokens and $0.50 per million output tokens, with a 50% discount for the first two weeks. The listing shows 96 upvotes at day rank 20 and an average of 4.7 out of 5 stars from seven reviews. The same page describes the GLM models as MIT-licensed.
What is not established
None of the three sources names the vendor behind the 100,000 chips. The claim that the run took place on home-grown hardware comes from the company and is relayed by SCMP; we found no independent confirmation of it. The sources use both the name Zhipu AI and the brand Z.ai for this same release.
FAQ
What is Ox Alpha?
The code name Zhipu AI used for GLM-5.3-Flash while the model was being tested ahead of its official announcement.
How much does the GLM-5.3-Flash API cost?
Product Hunt lists $0.15 per million input tokens and $0.50 per million output tokens, with a 50% discount for the first two weeks.
Where can I download GLM-5.3-Flash?
The weights are on Hugging Face under zai-org/GLM-5.3-Flash, and OpenRouter also hosts a free version of the model.