Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status

Qwen3.8-Flash

Budget

by Alibaba

Qwen3.8-Flash is the hosted build of Qwen3.8-Flash-Next, the open-weight model Alibaba published on August 26, 2026 as an explicit preview of the architecture that will underpin Qwen4. Keep the two straight: Flash-Next is the downloadable preview under a qwen-community-1.0 license with a 262,144 token native window and no list price, while Qwen3.8-Flash on QwenCloud is the managed product at $0.16 per million input tokens and $0.47 output with a 1 million token default context and official tool support. The architecture is the interesting part. It is 125 billion parameters with only 6 billion activated per token, paired with a 51 billion parameter n-gram embedding table indexing 20 million bigrams and trigrams, and a 4 billion parameter multi-token prediction layer. The n-gram table is the unusual lever: those parameters can sit in ordinary system RAM instead of GPU memory, which makes them far cheaper to hold than MoE experts. Alibaba reports better results than the 397 billion parameter Qwen3.7-Plus at roughly one ninth the training cost. On the self-reported table against DeepSeek-V4-Flash-0731 it leads DeepSWE 1.1 (58.7 to 54.4), SWE-bench Pro (62.5 to 56.0), Toolathlon Verified (73.5 to 70.3), GPQA Diamond (91.7 to 90.8), and LiveCodeBench v6 (91.9 to 90.6). NL2Repo-Bench is the honest miss at 48.1 against 54.2. Vision numbers are strong on paper (AndroidWorld 84.5, RealWorldQA 88.5, MathVision 90.6) but OSWorld 2.0 sits at 19.4, which is still a coin flip on full task success. Every figure here is vendor-reported.

Input Price

$0.16

per 1M tokens

Output Price

$0.47

per 1M tokens

Context Window

1M

tokens

Released

2026-08

API access

Capabilities

textvisionvideotool-usecodereasoning

Key Strengths

  • Only 6B active parameters per token out of 125B total
  • N-gram embedding table offloads 51B parameters to system RAM
  • 1M token default context on the hosted Cloud build
  • Leads DeepSeek V4 Flash on DeepSWE and SWE-bench Pro (vendor-reported)
  • Roughly one ninth the training cost of Qwen3.7-Plus
  • Open-weight Flash-Next preview available for self-hosting

Best For

  • High-volume agentic coding at flash pricing
  • Self-hosted multimodal agents on modest accelerator budgets
  • Long-context document and repository work
  • Teams evaluating the Qwen4 architecture ahead of the full release

Benchmark Scores

BenchmarkScoreDescription
GPQA Diamond91.7Graduate-level science questions verified by domain experts

Scores sourced from public benchmark datasets. See full benchmark leaderboard for all models.

Pricing Details

Input tokens

$0.16

per 1M tokens

Output tokens

$0.47

per 1M tokens

Estimated cost per 1K requests

$0.40

~1K input + ~500 output tokens avg

Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.

Related Models

View DocumentationCompare ModelsCost CalculatorFull Pricing Guide