Qwen3.8 Max Prime
Flagshipby Alibaba
Qwen3.8 Max Prime is Alibaba's own version of the speed-tier pattern Z.ai has been running on GLM-5.3: the same model, a faster lane, and a price that maps to the multiple. Released September 23, 2026, it serves identical weights to Qwen3.8-Max, the 2.4 trillion parameter mixture-of-experts activating roughly 95 billion per token, with the same 1 million token context window and native text, image, and video input. Alibaba lists it at $4 per million input tokens and $12 output, exactly double the $2 and $6 the standard endpoint charges, and its own documentation is explicit that model-supported capabilities and usage restrictions are the same as the original model. What the premium buys is throughput: Alibaba claims output-speed-sensitive scenarios see TPS raised to 1.5 to 2 times the standard API, with a soft ceiling rather than a hard rate limit, so a request is not throttled as long as the platform has spare resources. No controlled independent measurement of that speed claim has been published yet, and early OpenRouter traffic figures cited by third parties showed a far smaller gap (about 40 versus 37 tokens per second), so treat it as a vendor number until someone benchmarks the two lanes side by side. Every benchmark score published for Qwen3.8-Max, including the 86.6 on Terminal-Bench 2.1 and the split software-engineering picture against Fable 5 and GPT-5.6 Sol, applies unchanged to Prime, because nothing about the model itself moved.
Input Price
$4.00
per 1M tokens
Output Price
$12.00
per 1M tokens
Context Window
1M
tokens
Released
2026-09
API access
Capabilities
Key Strengths
- ✓Identical 2.4T MoE weights and benchmark profile to Qwen3.8-Max
- ✓Vendor-claimed 1.5 to 2x output throughput over the standard API
- ✓Soft ceiling rather than a hard rate limit when the platform has spare capacity
- ✓1M token context window with native text, image, and video input
- ✓Same tool-use and agentic capabilities as the base model
Best For
- ▸Output-speed-sensitive agent loops already built on Qwen3.8-Max
- ▸Interactive applications where streaming latency is the bottleneck
- ▸Teams that have hit rate limits on the standard Qwen3.8-Max endpoint
- ▸Workloads that can absorb 2x pricing for a throughput gain that has not yet been independently measured
Pricing Details
Input tokens
$4.00
per 1M tokens
Output tokens
$12.00
per 1M tokens
Estimated cost per 1K requests
$10.00
~1K input + ~500 output tokens avg
Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.
Related Models
Qwen3.8 27B
Mid-tierAlibaba
$0.50 in / $3.00 out
Qwen3.8-Omni-Flash
BudgetAlibaba
$0.15 in / $0.47 out
Qwen3.8-Flash
BudgetAlibaba
$0.15 in / $0.47 out
Qwen3.8 2.4T-A95B
FlagshipAlibaba
$2.00 in / $6.00 out
Qwen3.8-Max
FlagshipAlibaba
$2.00 in / $6.00 out
Qwen3.7-Max
FlagshipAlibaba
$2.50 in / $7.50 out
Claude Mythos 5.1
FlagshipAnthropic
$10.00 in / $50.00 out
Claude Opus 5.5
FlagshipAnthropic
$4.00 in / $20.00 out
Claude Opus 5
FlagshipAnthropic
$5.00 in / $25.00 out
Claude Fable 5.1
FlagshipAnthropic
$10.00 in / $50.00 out