Qwen/Qwen3.8-Flash-Next¶
Model Information¶
Qwen/Qwen3.8-Flash-Next is a high-throughput multimodal Mixture-of-Experts (MoE) reasoning model developed by Alibaba Cloud. It has 125 billion total parameters with 6 billion activated per token, and delivers 6–10× throughput compared to its predecessor at roughly 1/9 the training cost. It supports text and image input and is optimized for production-scale agentic workflows, long-context tasks, and office automation.
- Model Developer: Alibaba Cloud (Qwen Team)
- Model Release Date: August 26, 2026
- Supported Languages: 100+ languages (multilingual BPE tokenizer with ~151K base tokens).
- Applicable License: Qwen Community License 1.0
Model Architecture¶
Qwen/Qwen3.8-Flash-Next uses an experimental qwen4_exp architecture combining Gated DeltaNet linear-attention layers, Qwen Sparse Attention (QSA) layers, and a 51B N-gram Embedding lookup for zero-compute token disambiguation. This design enables a 7.6× prefill speedup and 4.9× decode speedup compared to full attention at 1M-token context.
Key Architecture Details:
- Model Type: Multimodal MoE (Mixture of Experts), image-text-to-text
- Parameters: 125B total, 6B activated per token; plus 51B N-gram embedding (zero-compute lookup) and 4B MTP head for speculative decoding
- MoE Configuration: 512 experts total; 10 routed + 1 shared activated per token (~2% activation rate)
- Architecture Layout: 48 layers in 12 macro-blocks: 3×(Gated DeltaNet → MoE) + 1×(QSA → MoE) per block
- Context Length: Up to 262,144 tokens natively; extensible to 1,000,000 tokens via YaRN
-
Training Strategy:
- Muon optimizer (2D linear maps) + AdamW (embeddings/router/low-rank params)
- Multi-Token Prediction (MTP) for speculative decoding
- Batch warmup eliminated, saving ~18.8% optimizer steps
- Training cost ~1/9 of Qwen3.7-Plus (vendor-reported)
-
Tokenizer: 248,320-token vocabulary with multilingual BPE
-
Capabilities:
- Agentic coding and repository-level code generation
- Office automation and multi-step agent workflows
- Vision and video understanding (multimodal input)
- Tool use, function calling, and multi-step planning
- Long-context production workloads (up to 1M tokens with YaRN)
- Thinking mode (reasoning_effort: xhigh / medium / low)
Benchmark Scores¶
All scores are self-reported by Alibaba Cloud at time of release; independent third-party verification is pending.
| Category | Benchmark | Score |
|---|---|---|
| Agentic Coding | SWE-bench Pro | 62.5% |
| SWE-bench Multilingual | 81.0% | |
| DeepSWE 1.1 | 58.7% | |
| Reasoning | GPQA Diamond | 91.7% |
| HLE | 35.9% | |
| LiveCodeBench v6 | 91.9% | |
| IFBench | 81.3% | |
| Agents / Office | CoWorkBench | 73.9% |
| JobBench | 55.7% | |
| Toolathlon Verified | 73.5% | |
| Multimodal | AndroidWorld | 84.5% |
| MathVision (CI) | 95.7% | |
| CharXiv RQ (CI) | 90.6% | |
| RealWorldQA | 88.5% |