Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status

GLM-5.3

Flagship

by Z.ai (Zhipu AI)

GLM-5.3 is Z.ai's August 2026 flagship, and the interesting part is how it was built: not a new pretrain, but extended post-training on the same base as GLM-5.2. The gains are concentrated where that kind of work pays off. Terminal-Bench 3.0 jumped from 4.6 to 28.3, DeepSWE v1.1 went from 46.2 to 66.9, and independent evaluation puts it at 95.4 on SWE-bench Verified. Z.ai also reports 62.5 on Humanity's Last Exam with tools and a notable cybersecurity capability that emerged rather than being targeted, with 54.4 on ExploitBench and 84.5 on CyberGym. Pricing is $1.40 per million input tokens and $4.40 output, with a 1M token context window and 128K max output. It launched API-first through the Z.ai API and Coding Plan; the Hugging Face repository is a gated placeholder counting down to August 28, 2026, so the weights are announced rather than available. Z.ai has not published a parameter count or a license for this release, and its predecessor's plain MIT terms should not be assumed to carry over. Treat every score here as vendor-reported except the independently run SWE-bench figure.

Input Price

$1.40

per 1M tokens

Output Price

$4.40

per 1M tokens

Context Window

1.0M

tokens

Released

2026-08

API access

Capabilities

texttool-usecodereasoning

Key Strengths

  • Terminal-Bench 3.0 of 28.3, up from 4.6 on GLM-5.2
  • 95.4 on SWE-bench Verified from an independent evaluator
  • 1M token context with 128K max output
  • $1.40/$4.40 pricing, far below comparable frontier coding models
  • Emergent cybersecurity capability: 84.5 CyberGym, 54.4 ExploitBench

Best For

  • Agentic coding and long-horizon software tasks
  • Security research and exploit reasoning
  • Cost-sensitive frontier workloads
  • Teams evaluating a self-hostable path once weights land

Benchmark Scores

BenchmarkScoreDescription
SWE-bench95.4Real-world software engineering tasks from GitHub issues (SWE-bench Verified)
MMLU-Pro86.8General knowledge and reasoning across 57 subjects
GPQA Diamond88.1Graduate-level science questions verified by domain experts
Humanity's Last Exam (tools)62.5Multidisciplinary expert-level reasoning with tool access

Scores sourced from public benchmark datasets. See full benchmark leaderboard for all models.

Pricing Details

Input tokens

$1.40

per 1M tokens

Output tokens

$4.40

per 1M tokens

Estimated cost per 1K requests

$3.60

~1K input + ~500 output tokens avg

Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.

Related Models

View DocumentationCompare ModelsCost CalculatorFull Pricing Guide