Qwen3.8-27B is open-source and ready to use! Ascend 0-Day adaptation: 27 billion parameter inference efficiency maximized.

Alibaba's Qwen team under the Qwen model family officially open-sourced the Qwen3.8-27B model on the evening of August 14. Ascend completed 0-Day adaptation immediately after the release, enabling efficient inference deployment of Qwen3.8-27B on Atlas 800 A3 and Atlas 850E supernodes via the vLLM-Ascend open-source inference engine.
Qwen3.8-27B is a natively multimodal dense model with 27 billion parameters. It natively supports a context length of 262K tokens, extendable to 1M tokens via YaRN technology.
Architecturally, Qwen3.8-27B adopts a hybrid linear attention design, with 48 of its 64 layers using Gated DeltaNet linear attention and 16 layers using Full Attention. This design breaks the O(n²) computational complexity for long sequences, balancing long-sequence processing efficiency with contextual modeling capability.
In terms of performance, Qwen3.8-27B significantly surpasses its predecessor Qwen3.6-27B in coding and office scenarios, with overall performance exceeding Qwen3.7-Plus.
The model scores 73.0 on the Agentic terminal coding test (previous generation: 63.4) and 61.7 on the SWE-bench Pro test (previous generation: 53.5). It also introduces a new reasoning_effort feature that dynamically adjusts thinking depth based on task difficulty to conserve computational resources.
Targeting the structural characteristics of Qwen3.8-27B, Ascend employs several key optimization techniques. In the Prefill phase, a GDN fusion operator processes long sequences efficiently in chunks, reducing intermediate data movement and operator scheduling overhead.
vLLM-Ascend supports W8A8 quantization deployment, quantizing weights and activations from BF16 to 8-bit to fully leverage the INT8/MXFP8 compute capabilities of Ascend NPUs.
Additionally, it leverages Qwen3.8's MTP speculative inference capability, where the MTP module predicts subsequent tokens while the main model performs parallel verification, reducing the number of main model invocations. The Decode phase supports capturing computation as device-side execution graphs (ACLGraph), eliminating repeated scheduling overhead.
Qwen3.8-27B is now available on Hugging Face and ModelScope simultaneously. Ascend has provided deployment guidance for the model, allowing users to quickly complete inference deployment through the vLLM-Ascend open-source inference engine.
Related Articles

Huawei phone users have finally got what they've been waiting for! The HarmonyOS trial beta version of NetEase Cloud Music is now available.
about 4 hours ago

The price after discounts is 6544 yuan! Lenovo ThinkPad E14 2026 laptop is now available: Core 5 320 + 512GB
about 5 hours ago


