Skip to content

Qwen/Qwen3.6-35B-A3B

Model Information

Qwen/Qwen3.6-35B-A3B is a multimodal Mixture-of-Experts (MoE) reasoning model developed by Alibaba Cloud. It has 35 billion total parameters with only 3 billion activated per token, delivering strong performance at low inference cost. It supports text and image input and excels at agentic coding, long-context processing, and multilingual tasks.

  • Model Developer: Alibaba Cloud (Qwen Team)
  • Model Release Date: April 16, 2026
  • Supported Languages: 100+ languages including English, Chinese, French, Spanish, German, Japanese, Korean, Portuguese, and others.
  • Applicable License: Apache 2.0

Model Architecture

Qwen/Qwen3.6-35B-A3B uses a hybrid MoE architecture combining Gated DeltaNet (linear attention) and Gated Attention layers with a vision encoder for multimodal input.

Key Architecture Details:

  • Model Type: Multimodal MoE (Mixture of Experts), image-text-to-text
  • Parameters: 35B total, 3B activated per token
  • MoE Configuration: 256 experts total; 8 routed + 1 shared activated per token
  • Context Length: Up to 262,144 tokens natively; extensible to 1,010,000 tokens via YaRN
  • Max Output: 80,000 tokens
  • Training Strategy:

    • Pre-training and post-training with Multi-Token Prediction (MTP)
    • Optimized for repository-level and agentic coding tasks
    • Thinking mode with preserve_thinking support
  • Tokenizer: 248,320-token vocabulary with multilingual BPE

  • Capabilities:

    • Agentic and repository-level code generation
    • Vision and video understanding (multimodal input)
    • Tool use and function calling
    • Long-context processing (up to 1M tokens with YaRN)
    • Multilingual reasoning

Benchmark Scores

All scores are self-reported by Alibaba Cloud at time of release.

Category Benchmark Score
Agentic Coding SWE-bench Verified 73.4%
SWE-bench Pro 49.5%
SWE-bench Multilingual 67.2%
Reasoning GPQA Diamond 86.0%
AIME 2026 92.7%
General MMLU-Pro 85.2%
LiveCodeBench 80.4%
Multimodal MMMU-Pro 75.0%
RealWorldQA 85.3%
OmniDocBench 89.9%
Long Context Context Arena 83.5%

References