Qwen3.8-Flash
Budgetby Alibaba
Qwen3.8-Flash is the hosted build of Qwen3.8-Flash-Next, the open-weight model Alibaba published on August 26, 2026 as an explicit preview of the architecture that will underpin Qwen4. Keep the two straight: Flash-Next is the downloadable preview under a qwen-community-1.0 license with a 262,144 token native window and no list price, while Qwen3.8-Flash on QwenCloud is the managed product at $0.16 per million input tokens and $0.47 output with a 1 million token default context and official tool support. The architecture is the interesting part. It is 125 billion parameters with only 6 billion activated per token, paired with a 51 billion parameter n-gram embedding table indexing 20 million bigrams and trigrams, and a 4 billion parameter multi-token prediction layer. The n-gram table is the unusual lever: those parameters can sit in ordinary system RAM instead of GPU memory, which makes them far cheaper to hold than MoE experts. Alibaba reports better results than the 397 billion parameter Qwen3.7-Plus at roughly one ninth the training cost. On the self-reported table against DeepSeek-V4-Flash-0731 it leads DeepSWE 1.1 (58.7 to 54.4), SWE-bench Pro (62.5 to 56.0), Toolathlon Verified (73.5 to 70.3), GPQA Diamond (91.7 to 90.8), and LiveCodeBench v6 (91.9 to 90.6). NL2Repo-Bench is the honest miss at 48.1 against 54.2. Vision numbers are strong on paper (AndroidWorld 84.5, RealWorldQA 88.5, MathVision 90.6) but OSWorld 2.0 sits at 19.4, which is still a coin flip on full task success. Every figure here is vendor-reported.
Input Price
$0.16
per 1M tokens
Output Price
$0.47
per 1M tokens
Context Window
1M
tokens
Released
2026-08
API access
Capabilities
Key Strengths
- ✓Only 6B active parameters per token out of 125B total
- ✓N-gram embedding table offloads 51B parameters to system RAM
- ✓1M token default context on the hosted Cloud build
- ✓Leads DeepSeek V4 Flash on DeepSWE and SWE-bench Pro (vendor-reported)
- ✓Roughly one ninth the training cost of Qwen3.7-Plus
- ✓Open-weight Flash-Next preview available for self-hosting
Best For
- ▸High-volume agentic coding at flash pricing
- ▸Self-hosted multimodal agents on modest accelerator budgets
- ▸Long-context document and repository work
- ▸Teams evaluating the Qwen4 architecture ahead of the full release
Benchmark Scores
| Benchmark | Score | Description |
|---|---|---|
| GPQA Diamond | 91.7 | Graduate-level science questions verified by domain experts |
Scores sourced from public benchmark datasets. See full benchmark leaderboard for all models.
Pricing Details
Input tokens
$0.16
per 1M tokens
Output tokens
$0.47
per 1M tokens
Estimated cost per 1K requests
$0.40
~1K input + ~500 output tokens avg
Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.
Related Models
Qwen3.8 27B
Mid-tierAlibaba
$0.42 in / $2.55 out
Qwen3.8 2.4T-A95B
FlagshipAlibaba
$2.00 in / $6.00 out
Qwen3.8-Max
FlagshipAlibaba
$2.00 in / $6.00 out
Qwen3.7-Max
FlagshipAlibaba
$2.50 in / $7.50 out
Claude Haiku 4.5
BudgetAnthropic
$1.00 in / $5.00 out
GPT-5.6 Luna
BudgetOpenAI
$0.20 in / $1.20 out
GPT-4o-mini
BudgetOpenAI
$0.15 in / $0.60 out
Gemini 2.0 Flash
Budget$0.10 in / $0.40 out