Qwen3.8-Max
Flagshipby Alibaba
Qwen3.8-Max is Alibaba's flagship closed model, released August 3, 2026 at $2 per million input tokens and $6 per million output, a cut from the $2.50/$7.50 charged for Qwen3.7-Max. It is a Mixture-of-Experts architecture with roughly 2.4 trillion total parameters and about 95 billion active per token, a 1 million token context window, and native text, image, and video input. Caching is aggressive: implicit cache reads run $0.25 per million tokens, explicit cache creation $2.50, and explicit cache reads $0.17, which materially changes the economics of repeated long prompts. The benchmark picture is genuinely split rather than uniformly strong. Qwen reports 86.6 on Terminal-Bench 2.1, ahead of Claude Opus 4.8 and Claude Fable 5 at 84.6 but behind GPT-5.6 Sol at 88.8, and it leads PaperBench at 93.0 and IFBench at 82.8. On core software engineering it is clearly behind the Western frontier: 67.7 against Fable 5's 80.0 on SWE-bench Pro, and 73.5 against 88.8 on FrontierSWE. Read it as a strong agentic and instruction-following model at roughly a fifth of frontier flagship pricing, not a SWE-bench leader. Alibaba paired the release with Qwen3.8-27B open weights on August 14, giving the family a permissive local option alongside the hosted flagship. Available through Alibaba Cloud Model Studio and the Qwen API.
Input Price
$2.00
per 1M tokens
Output Price
$6.00
per 1M tokens
Context Window
1M
tokens
Released
2026-08
API access
Capabilities
Key Strengths
- ✓$2/$6 per 1M tokens, down from $2.50/$7.50 on Qwen3.7-Max
- ✓2.4T parameter MoE with roughly 95B active per token
- ✓1M token context window with native text, image, and video input
- ✓86.6 on Terminal-Bench 2.1, ahead of Opus 4.8 and Fable 5
- ✓Leads PaperBench at 93.0 and IFBench at 82.8
- ✓Cheap cache reads at $0.25 implicit and $0.17 explicit per 1M
Best For
- ▸Agentic tool use and long-horizon task execution
- ▸Multimodal document, chart, and video understanding
- ▸Instruction-heavy production pipelines at flagship-adjacent quality
- ▸Cost-sensitive workloads that would otherwise use a frontier flagship
Pricing Details
Input tokens
$2.00
per 1M tokens
Output tokens
$6.00
per 1M tokens
Estimated cost per 1K requests
$5.00
~1K input + ~500 output tokens avg
Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.