GLM-5.3-FlashX
BudgetGLM-5.3-FlashX is the clearest example yet of a price that buys latency rather than capability. Released September 18, 2026, it serves the same weights as GLM-5.3 Flash, the same 320 billion parameter mixture-of-experts activating roughly 18 billion per token, on a faster stack. Z.ai lists it at $0.37 per million input tokens, $0.075 cached, and $1.25 output, which is roughly two and a half times the $0.15 and $0.50 the base Flash model charges, with cache storage free for a limited time. What you get for the premium is throughput: Z.ai advertises up to 200 output tokens per second, and independent measurement on its own API has landed nearer 98. Context is 1 million tokens with up to 131,072 output, and the modality set carries over unchanged, taking text, images, video, and files. The decision rule is simple and unusually clean, because capability is not part of it. If a user is waiting on the stream, FlashX is worth the multiple; if a batch job is running overnight, it is not. Note that FlashX is a hosted tier only and has no separate weight release: the MIT-licensed checkpoint on Hugging Face is GLM-5.3-Flash, and self-hosting it gives you the model but not the serving stack that FlashX is selling.
Input Price
$0.37
per 1M tokens
Output Price
$1.25
per 1M tokens
Context Window
1.0M
tokens
Released
2026-09
API access
Capabilities
Key Strengths
- ✓Up to 200 output tokens per second, vendor-reported
- ✓Identical 320B MoE weights to GLM-5.3 Flash, so capability is unchanged
- ✓$0.37/$1.25 with cached input at $0.075
- ✓1M token context with 131K max output
- ✓Text, image, video, and file input
- ✓Cache storage free for a limited time
Best For
- ▸Interactive coding and chat where a user waits on the stream
- ▸Latency-sensitive agent loops
- ▸Computer-use and visual understanding at speed
- ▸Teams already on GLM-5.3 Flash that need a faster tier without a model change
Pricing Details
Input tokens
$0.37
per 1M tokens
Output tokens
$1.25
per 1M tokens
Estimated cost per 1K requests
$0.99
~1K input + ~500 output tokens avg
Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.
Related Models
GLM-5.3
FlagshipZ.ai (Zhipu AI)
$1.40 in / $4.40 out
GLM-5.3 Flash
BudgetZ.ai (Zhipu AI)
$0.15 in / $0.50 out
GLM-5.2
FlagshipZ.ai (Zhipu AI)
$1.40 in / $4.40 out
Claude Haiku 4.5
BudgetAnthropic
$1.00 in / $5.00 out
GPT-6 Luna
BudgetOpenAI
$0.10 in / $0.50 out
GPT-5.6 Luna
BudgetOpenAI
$0.20 in / $1.20 out
GPT-4o-mini
BudgetOpenAI
$0.15 in / $0.60 out