Which Chinese large language model is the strongest? Qwen ranks first, DeepSeek didn't make the top three.

SuperCLUE released the complete Chinese large model benchmark rankings for July 2026, with a total of 12 mainstream domestic models participating in the tests.
In the overall score ranking, Qwen3.8-Max took first place, Kimi-K3 ranked second, Doubao-Seed-2.1-pro came in third, and DeepSeek-V4-Flash placed fourth.
This evaluation differs from previous ones, upgrading traditional code generation to agentic programming with a direct weight of 25%, better aligning with programmers' real development workflows.
The remaining five dimensions are mathematical reasoning at 18%, scientific reasoning at 16%, precise instruction following at 16%, hallucination control at 16%, and agent task planning at 9%, with the final score calculated across all six dimensions.
In terms of overall scores, Qwen3.8-Max scored 71.48 points and Kimi-K3 scored 70.68 points, a gap of just 0.8 points.
Doubao-Seed-2.1-pro scored 65.95 points, DeepSeek-V4-Flash scored 65.6 points, and DeepSeek-V4-Pro ranked fifth with 64.4 points.
The strengths and weaknesses across different segments vary significantly. Kimi-K3 ranked first overall in agentic programming with 75.79 points, showing clear advantages in writing engineering code and multi-round development interactions; Qwen ranked second with 68.42 points. DeepSeek-V4-Flash was strongest in mathematical reasoning with 78.95 points, while Doubao performed best in scientific reasoning with 77.19 points.
In the two categories of hallucination control and agent task planning, Qwen led by a decisive margin, being less likely to fabricate false information and more capable of breaking down complex tasks.
Inference speed and cost-effectiveness are another set of criteria. The entire DeepSeek-V4 series falls into the high cost-performance tier, with low API pricing and reliable inference quality. Meanwhile, DeepSeek-V4-Flash, GLM-5.2, and Hy3 sit in the high-efficiency zone with short computation times.
By contrast, Qwen and Kimi, the top two in overall scores, have high reasoning scores but long per-inference wait times, placing them in the low-efficiency tier. Kimi's latency is even three times that of the fastest model.
Developers with limited budgets handling batch development should prioritize the DeepSeek series; for daily office work, copywriting, and concerns about AI fabrication, Qwen is more suitable; for professional programmers writing engineering code over the long term, Kimi provides a better experience.
Related Articles

Huawei phone users have finally got what they've been waiting for! The HarmonyOS trial beta version of NetEase Cloud Music is now available.
about 4 hours ago

The price after discounts is 6544 yuan! Lenovo ThinkPad E14 2026 laptop is now available: Core 5 320 + 512GB
about 5 hours ago



