Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status

Qwen3.8-Max

Flagship

by Alibaba

Qwen3.8-Max is Alibaba's flagship closed model, released August 3, 2026 at $2 per million input tokens and $6 per million output, a cut from the $2.50/$7.50 charged for Qwen3.7-Max. It is a Mixture-of-Experts architecture with roughly 2.4 trillion total parameters and about 95 billion active per token, a 1 million token context window, and native text, image, and video input. Caching is aggressive: implicit cache reads run $0.25 per million tokens, explicit cache creation $2.50, and explicit cache reads $0.17, which materially changes the economics of repeated long prompts. The benchmark picture is genuinely split rather than uniformly strong. Qwen reports 86.6 on Terminal-Bench 2.1, ahead of Claude Opus 4.8 and Claude Fable 5 at 84.6 but behind GPT-5.6 Sol at 88.8, and it leads PaperBench at 93.0 and IFBench at 82.8. On core software engineering it is clearly behind the Western frontier: 67.7 against Fable 5's 80.0 on SWE-bench Pro, and 73.5 against 88.8 on FrontierSWE. Read it as a strong agentic and instruction-following model at roughly a fifth of frontier flagship pricing, not a SWE-bench leader. Alibaba paired the release with Qwen3.8-27B open weights on August 14, giving the family a permissive local option alongside the hosted flagship. Available through Alibaba Cloud Model Studio and the Qwen API.

Input Price

$2.00

per 1M tokens

Output Price

$6.00

per 1M tokens

Context Window

1M

tokens

Released

2026-08

API access

Capabilities

textvisionvideocodereasoningtool-use

Key Strengths

  • $2/$6 per 1M tokens, down from $2.50/$7.50 on Qwen3.7-Max
  • 2.4T parameter MoE with roughly 95B active per token
  • 1M token context window with native text, image, and video input
  • 86.6 on Terminal-Bench 2.1, ahead of Opus 4.8 and Fable 5
  • Leads PaperBench at 93.0 and IFBench at 82.8
  • Cheap cache reads at $0.25 implicit and $0.17 explicit per 1M

Best For

  • Agentic tool use and long-horizon task execution
  • Multimodal document, chart, and video understanding
  • Instruction-heavy production pipelines at flagship-adjacent quality
  • Cost-sensitive workloads that would otherwise use a frontier flagship

Pricing Details

Input tokens

$2.00

per 1M tokens

Output tokens

$6.00

per 1M tokens

Estimated cost per 1K requests

$5.00

~1K input + ~500 output tokens avg

Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.

Related Models

View DocumentationCompare ModelsCost CalculatorFull Pricing Guide