100,000 domestic computing cards power "Niu Lai": Zhipu open-sources GLM-5.3-Flash

The answer arrived as scheduled. On the evening of August 26, Zhipu announced the launch and open-sourcing of GLM-5.3-Flash (320B-A18B), officially confirming it was the anonymous model Ox-Alpha that had been trending in overseas developer communities over the past week—known in Chinese communities as "Niu Lai."
This model is the first native multimodal model in the GLM-5 series, with 320B total parameters and only 18B activated parameters. It adopts the MIT license for open weights and was listed on Hugging Face immediately upon release.
On August 20, Ox-Alpha quietly launched as a "stealth model" on OpenRouter, the world's largest model aggregation platform. It topped daily call volume on its first day, setting a single-day historical record, while also ending DeepSeek's 56-day reign atop OpenCode. It processed 62T tokens in a single day, quickly becoming the most popular model on both platforms that week and breaking historical records on both.
During testing, an independent researcher published a technical fingerprint report stating that its tokenizer produced token counts identical to GLM-5.3 across 25 prompt sets, and that 11 probe cross-tests fully matched the GLM family, pointing to Zhipu with "99% confidence"—now officially confirmed.
According to reports, the model's capabilities directly benchmark against international flagships. GLM-5.3-Flash scored 57 on the Artificial Analysis Intelligence Index, tying with Anthropic's most popular Claude Opus 4.8, and surpassing GLM-5.2's 53 and DeepSeek V4 Pro's official 53.
Pricing breaks new ground. GLM-5.3-Flash is priced at 1/10 of GLM-5.3, dropping to as low as 1/20 during the limited-time discount period, and just 1/40 of Opus 4.8. Specifically, the model costs 0.8 yuan per million input tokens, 2.8 yuan per million output tokens, and 0.23 yuan for cache hits.
Domestic computing power goes global at scale for the first time. Zhipu disclosed that all online traffic during "Niu Lai's" testing phase, as well as the online services following GLM-5.3-Flash's release, were powered by 100,000 domestic chips.
This marks the first time domestic computing chips have directly served real global workloads at scale. The market debate over "where the daily 100 trillion token computing supply comes from" now has its answer.
Additionally, a reporter from the STAR Market Daily learned from relevant sources that these domestic computing chips may come from Huawei, Hygon, or Moore Threads.
According to Zhipu's introduction to the STAR Market Daily reporter, the technical architecture is the root of cost reduction. GLM-5.3-Flash was pre-trained on 30T tokens of multimodal data, marking the first time the GLM series introduced a hybrid architecture of sparse attention and linear attention, along with the innovative manifold-constrained hyper-connection (mHC).
Per Zhipu's disclosure, this architecture reduces attention computation by approximately 3x compared to GLM-5.3, shrinks KV cache by 4.4x, and supports cost-effective deployment of a 1M context window.
Open weights are immediately available. GLM-5.3-Flash is now fully open-sourced globally, integrated into coding platforms like ZCode, and included in the GLM Coding Plan (with 10,000 trial cards distributed daily), with API access opened simultaneously.
Zhang Youyu, Secretary-General of the AI Committee under the Beijing Computer Society and Distinguished Researcher at Peking University, told the STAR Market Daily reporter that the strategy of "anonymous launch, public testing, and scheduled reveal" allowed GLM-5.3-Flash to complete real-traffic stress testing and capture blind-test reputation before its unveiling;
Against the backdrop of DeepSeek V4 Flash raising prices and growing market demand for "affordable frontier models," it enters with the combination of "frontier capability + 1/40 price + domestic computing infrastructure," once again breaking through the price floor for frontier models.
Related Articles

Huawei phone users have finally got what they've been waiting for! The HarmonyOS trial beta version of NetEase Cloud Music is now available.
about 3 hours ago

The price after discounts is 6544 yuan! Lenovo ThinkPad E14 2026 laptop is now available: Core 5 320 + 512GB
about 4 hours ago


