Qwen3.8-Flash open-sourced, surpassing V4 Flash: early adoption of next-generation Qwen4 architecture

Tonight, in addition to Zhipu open-sourcing the Niu Lai large model, Alibaba's Qwen also released the Qwen3.8-Flash large model. Although it still uses the 3.X version name, it is actually based on the next-generation Qwen4 architecture.
Qwen previously explained that Qwen3.8-Flash is designed to let the open-source community get an early experience of the new technologies in the Qwen4 architecture, serving as a technical preview of the next-generation model.
According to official information, Qwen3.8-Flash is a multimodal MoE model with 125B parameters plus 51B N-gram embeddings, activating only 6B per token, with a 262K native context that can be extended to 1M via YaRN.
The company states it offers unmatched cost-effectiveness, priced at just $0.16 per 1M input tokens and $0.47 per 1M output tokens.
For comparison, DS V4 Flash currently costs $0.22 and $0.66 respectively, and prices double during peak hours, so Qwen3.8-Flash is indeed highly cost-effective.
Technically, Qwen3.8-Flash adopts the next-generation architecture—GDN + QSA hybrid attention, gated residual connections, N-gram embeddings, and the Muon optimizer—serving as the precursor to the architecture used in Qwen4.
In terms of performance benchmarks, Qwen3.8-Flash scores 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI).
Overall, this performance surpasses V4 Flash, and with its high cost-effectiveness, if real-world performance matches these numbers, it might be worth switching to this large model.
Related Articles

Huawei phone users have finally got what they've been waiting for! The HarmonyOS trial beta version of NetEase Cloud Music is now available.
about 3 hours ago

The price after discounts is 6544 yuan! Lenovo ThinkPad E14 2026 laptop is now available: Core 5 320 + 512GB
about 4 hours ago


