GLM-5.3
FlagshipGLM-5.3 is Z.ai's August 2026 flagship, and the interesting part is how it was built: not a new pretrain, but extended post-training on the same base as GLM-5.2. The gains are concentrated where that kind of work pays off. Terminal-Bench 3.0 jumped from 4.6 to 28.3, DeepSWE v1.1 went from 46.2 to 66.9, and independent evaluation puts it at 95.4 on SWE-bench Verified. Z.ai also reports 62.5 on Humanity's Last Exam with tools and a notable cybersecurity capability that emerged rather than being targeted, with 54.4 on ExploitBench and 84.5 on CyberGym. Pricing is $1.40 per million input tokens and $4.40 output, with a 1M token context window and 128K max output. It launched API-first through the Z.ai API and Coding Plan; the Hugging Face repository is a gated placeholder counting down to August 28, 2026, so the weights are announced rather than available. Z.ai has not published a parameter count or a license for this release, and its predecessor's plain MIT terms should not be assumed to carry over. Treat every score here as vendor-reported except the independently run SWE-bench figure.
Input Price
$1.40
per 1M tokens
Output Price
$4.40
per 1M tokens
Context Window
1.0M
tokens
Released
2026-08
API access
Capabilities
Key Strengths
- ✓Terminal-Bench 3.0 of 28.3, up from 4.6 on GLM-5.2
- ✓95.4 on SWE-bench Verified from an independent evaluator
- ✓1M token context with 128K max output
- ✓$1.40/$4.40 pricing, far below comparable frontier coding models
- ✓Emergent cybersecurity capability: 84.5 CyberGym, 54.4 ExploitBench
Best For
- ▸Agentic coding and long-horizon software tasks
- ▸Security research and exploit reasoning
- ▸Cost-sensitive frontier workloads
- ▸Teams evaluating a self-hostable path once weights land
Benchmark Scores
| Benchmark | Score | Description |
|---|---|---|
| SWE-bench | 95.4 | Real-world software engineering tasks from GitHub issues (SWE-bench Verified) |
| MMLU-Pro | 86.8 | General knowledge and reasoning across 57 subjects |
| GPQA Diamond | 88.1 | Graduate-level science questions verified by domain experts |
| Humanity's Last Exam (tools) | 62.5 | Multidisciplinary expert-level reasoning with tool access |
Scores sourced from public benchmark datasets. See full benchmark leaderboard for all models.
Pricing Details
Input tokens
$1.40
per 1M tokens
Output tokens
$4.40
per 1M tokens
Estimated cost per 1K requests
$3.60
~1K input + ~500 output tokens avg
Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.