100,000 domestic cards supported! Zhipu releases lightweight flagship model to rival DeepSeek

Zhipu AI has officially launched the lightweight flagship model GLM-5.3 Flash, positioned to compete directly with DeepSeek V4 Flash. Its online inference services are powered by a cluster of over 100,000 domestic chips, ensuring the entire pipeline runs on domestic hardware.
The model adopts a MoE (Mixture of Experts) architecture with 320B total parameters and only 18B activated parameters, placing it in the lightweight flagship category. It natively supports multimodal input, handling images, text, and video content. Unlike conventional small models distilled from larger ones, GLM-5.3 Flash features a brand-new independent architecture design rather than a simple pruned version of a large model. Its comprehensive evaluation scores surpass the previous flagship GLM-5.2, matching the level of mainstream international closed-source models.
In terms of pricing, GLM-5.3 Flash is priced at just one-tenth of GLM-5.3, with a 50% discount for the first two weeks after launch, further reducing invocation costs. Compared to DeepSeek V4 Flash, it offers cost advantages in most routine business scenarios.
Previously, this model was anonymously tested on overseas platforms under the codename Ox Alpha, accumulating over 50TB of traffic in five days and setting a new platform traffic growth record, validating market demand.
On the computing power front, this 100,000-scale domestic chip inference cluster is compatible with multiple domestic hardware vendors. Zhipu has not disclosed specific chip models, but industry information indicates the cluster's computing power comes from domestic chip manufacturers such as Huawei Ascend, Hygon, and Moore Threads.
This also means the model's online services no longer heavily rely on overseas GPUs, validating the feasibility of large-scale deployment of domestic large models combined with domestic computing power.
Industry experts believe that as enterprise AI invocation volumes surge, lightweight flagship models balancing performance and low cost are becoming the market mainstream. The release of GLM-5.3 Flash, on one hand, directly competes with DeepSeek's same-tier products; on the other hand, running large model inference on a 100,000-scale domestic chip cluster marks a significant step forward for domestic computing infrastructure.
Image source: Zhipu AI official website
Related Articles

Huawei phone users have finally got what they've been waiting for! The HarmonyOS trial beta version of NetEase Cloud Music is now available.
about 3 hours ago

The price after discounts is 6544 yuan! Lenovo ThinkPad E14 2026 laptop is now available: Core 5 320 + 512GB
about 4 hours ago


