GLM-5.3 Flash
BudgetGLM-5.3 Flash is not a cheaper GLM-5.3. It is a different model, released August 26, 2026, and the first natively multimodal member of the GLM-5 family: text, image, and video in one stack rather than a vision head bolted onto a text model. Architecturally it is a 320 billion parameter mixture-of-experts activating roughly 18 billion per token across 45 layers, with hybrid linear and sparse attention (IndexPool and mHC) and a 30 trillion token multimodal pretrain. Z.ai cites roughly 3.0x less attention compute and a 4.4x smaller KV cache than text GLM-5.3. Context is 1 million tokens, weights are MIT licensed on Hugging Face at zai-org/GLM-5.3-Flash, and the model ran on OpenCode and OpenRouter under the stealth id ox-alpha before launch. Pricing is the headline and also the trap: list is $0.15 per million input tokens, $0.03 cached, and $0.50 output, roughly a tenth of the $1.40 input rate on text GLM-5.3, but a 50 percent launch promotion runs at $0.075/$0.25 through September 9, 2026 at 16:00 UTC. Budget the list rate. On the self-reported table the movement is concentrated in agentic work: DeepSWE 63.4 against GLM-5.2 at 46.2, AutomationBench 48.8 against 26.2, Toolathlon 78.4 against 59.9. Terminal Bench 2.1 at 84.3 is a 3.3 point gain and HLE with tools at 55.3 against 54.7 is a wash. On vision it leads DeepSeek V4 Flash Vision Exp on Chartography (78.0 to 64.3) and OfficeQA-Pro (62.4 to 57.9), and it loses BabyVision to Gemini 3.7 Flash 53.4 to 70.9. One operational constraint: thinking cannot be turned off, since thinking.type accepts enabled only.
Input Price
$0.15
per 1M tokens
Output Price
$0.50
per 1M tokens
Context Window
1.0M
tokens
Released
2026-08
Open source
Capabilities
Key Strengths
- ✓First natively multimodal GLM-5: text, image, and video input
- ✓MIT weights at 18B active parameters with a 1M context window
- ✓List $0.15/$0.50, roughly a tenth of text GLM-5.3 on input
- ✓DeepSWE 63.4 and AutomationBench 48.8, up 17.2 and 22.6 on GLM-5.2
- ✓Roughly 3.0x less attention compute and a 4.4x smaller KV cache than GLM-5.3
- ✓3x the usable Coding Plan quota versus GLM-5.3
Best For
- ▸Multimodal coding and computer-use loops at flash pricing
- ▸Chart, document, and video understanding at high volume
- ▸Self-hosted multimodal agents on MIT terms
- ▸Long-context agentic pipelines where GLM-5.3 is too expensive per call
Benchmark Scores
| Benchmark | Score | Description |
|---|---|---|
| Humanity's Last Exam (tools) | 55.3 | Multidisciplinary expert-level reasoning with tool access |
Scores sourced from public benchmark datasets. See full benchmark leaderboard for all models.
Pricing Details
Input tokens
$0.15
per 1M tokens
Output tokens
$0.50
per 1M tokens
Estimated cost per 1K requests
$0.40
~1K input + ~500 output tokens avg
Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.
Open Source Model
GLM-5.3 Flash is free to download and self-host under the MIT. Hosted API pricing varies by provider (e.g., Together, Fireworks, Groq). See our open source LLM guide for deployment options.