Qwen/Qwen3.6-35B-A3B¶
Model Information¶
Qwen/Qwen3.6-35B-A3B is a multimodal Mixture-of-Experts (MoE) reasoning model developed by Alibaba Cloud. It has 35 billion total parameters with only 3 billion activated per token, delivering strong performance at low inference cost. It supports text and image input and excels at agentic coding, long-context processing, and multilingual tasks.
- Model Developer: Alibaba Cloud (Qwen Team)
- Model Release Date: April 16, 2026
- Supported Languages: 100+ languages including English, Chinese, French, Spanish, German, Japanese, Korean, Portuguese, and others.
- Applicable License: Apache 2.0
Model Architecture¶
Qwen/Qwen3.6-35B-A3B uses a hybrid MoE architecture combining Gated DeltaNet (linear attention) and Gated Attention layers with a vision encoder for multimodal input.
Key Architecture Details:
- Model Type: Multimodal MoE (Mixture of Experts), image-text-to-text
- Parameters: 35B total, 3B activated per token
- MoE Configuration: 256 experts total; 8 routed + 1 shared activated per token
- Context Length: Up to 262,144 tokens natively; extensible to 1,010,000 tokens via YaRN
- Max Output: 80,000 tokens
-
Training Strategy:
- Pre-training and post-training with Multi-Token Prediction (MTP)
- Optimized for repository-level and agentic coding tasks
- Thinking mode with
preserve_thinkingsupport
-
Tokenizer: 248,320-token vocabulary with multilingual BPE
-
Capabilities:
- Agentic and repository-level code generation
- Vision and video understanding (multimodal input)
- Tool use and function calling
- Long-context processing (up to 1M tokens with YaRN)
- Multilingual reasoning
Benchmark Scores¶
All scores are self-reported by Alibaba Cloud at time of release.
| Category | Benchmark | Score |
|---|---|---|
| Agentic Coding | SWE-bench Verified | 73.4% |
| SWE-bench Pro | 49.5% | |
| SWE-bench Multilingual | 67.2% | |
| Reasoning | GPQA Diamond | 86.0% |
| AIME 2026 | 92.7% | |
| General | MMLU-Pro | 85.2% |
| LiveCodeBench | 80.4% | |
| Multimodal | MMMU-Pro | 75.0% |
| RealWorldQA | 85.3% | |
| OmniDocBench | 89.9% | |
| Long Context | Context Arena | 83.5% |