LIVE
+++ Ox Alpha Revealed as GLM-5.3-Flash: Zhipu Stock Up 12% +++ Nvidia Nears $12.9 Billion Deal for Hugging Face +++ Who Gains as OpenAI's Leadership Ranks Empty Out +++ Hugging Face's $399 Microduck robot waddles and skates +++ OpenAI Report: 688 Agents Attacked Hugging Face +++ Gemini Omni 1.1 Flash: 40-second scenes, 360p drafts ++++++ Ox Alpha Revealed as GLM-5.3-Flash: Zhipu Stock Up 12% +++ Nvidia Nears $12.9 Billion Deal for Hugging Face +++ Who Gains as OpenAI's Leadership Ranks Empty Out +++ Hugging Face's $399 Microduck robot waddles and skates +++ OpenAI Report: 688 Agents Attacked Hugging Face +++ Gemini Omni 1.1 Flash: 40-second scenes, 360p drafts +++
All news ›
AI IN LIFE AI IN LIFENEWS
DAILY
DE EN
MODELS

Ox Alpha Revealed as GLM-5.3-Flash: Zhipu Stock Up 12%

The mystery Ox Alpha model is GLM-5.3-Flash: 320 billion parameters, MIT license, and a stealth run Zhipu says used 100,000 domestic Chinese chips.

Ox Alpha Revealed as GLM-5.3-Flash: Zhipu Stock Up 12%

Illustration: A sprawling data center hall lined with long rows of dark server racks under cold blue light.

Ox Alpha is GLM-5.3-Flash, an open-weight mixture-of-experts model from Zhipu AI whose reveal lifted the company's Hong Kong-listed shares by at least 12 percent.

At a glance

  • Ox Alpha is officially GLM-5.3-Flash: 320 billion total parameters, 18 billion active per query, mixture-of-experts design.
  • The stealth trial ran on a cluster of 100,000 chips made in China, per SCMP; the chip supplier is not named.
  • Zhipu shares closed at least 12 percent higher in Hong Kong on Thursday at HK$1,160.
  • Weights are published under an MIT license on Hugging Face (zai-org/GLM-5.3-Flash).
  • List price: $0.15 per million input tokens and $0.50 per million output tokens, with 50 percent off for two weeks.

Zhipu AI has confirmed that Ox Alpha, the anonymous system that climbed coding leaderboards for weeks, is GLM-5.3-Flash, and has released the model's weights. Shares closed at least 12 percent higher in Hong Kong on Thursday at HK$1,160. The detail drawing the most attention is not the model but where it ran: a cluster of 100,000 chips manufactured in China, according to the South China Morning Post.

Inside the model

GLM-5.3-Flash is a mixture-of-experts system with 320 billion total parameters and 18 billion active on any given query. SiliconANGLE reports an input context of 1 million tokens covering text, images and video, output of up to 131,072 tokens, and training on 30 trillion tokens. Z.ai says the model is ten times more cost-efficient than its predecessor.

On benchmarks, the company claims the top score on GDPval-AA v2 and second place on AutomationBench. Those are vendor figures, and no independent replication has been published so far.

The hardware is the real story

Before the formal launch, the model processed 62 trillion tokens during its stealth run, the SCMP reports. In the week of the reveal it accounted for 10.3 trillion tokens on OpenRouter, roughly 31 percent of that platform's weekly volume, and ranked first among coding systems there.

That matters well beyond the leaderboard. Serving inference at global scale on domestic silicon is the sharpest test yet of whether Chinese developers can operate without Nvidia accelerators while US export controls hold. The report stops short of naming the company that supplied the 100,000 chips.

License, price and availability

The weights are posted on Hugging Face as zai-org/GLM-5.3-Flash, and the GLM line carries an MIT license. Z.ai's own product listing puts pricing at $0.15 per million input tokens and $0.50 per million output tokens, with a 50 percent discount for the first two weeks. OpenRouter is hosting a free version at launch.

What is still unverified

Three claims remain open in the available reporting: the identity of the chip vendor, whether training as well as inference ran on domestic hardware, and any outside check of the benchmark results. Until those are settled, the performance numbers are best read as vendor claims relayed by the outlets cited below.

◈ AI-GENERATED REPORT · SOURCES LINKED

FAQ

What is Ox Alpha?

Ox Alpha was the code name Zhipu AI used to test its new model anonymously. It is GLM-5.3-Flash, and the weights are now available on Hugging Face.

Does GLM-5.3-Flash actually run on Chinese chips?

The SCMP reports that the stealth deployment ran on a cluster of 100,000 chips made in China. The chip maker is not identified in the report, and whether training also ran on that hardware is unclear.

How much does GLM-5.3-Flash cost?

Z.ai's product listing shows $0.15 per million input tokens and $0.50 per million output tokens, discounted by 50 percent for the first two weeks. OpenRouter offers a free hosted version at launch.