Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status

GLM-5.3-FlashX

Budget

by Z.ai (Zhipu AI)

GLM-5.3-FlashX is the clearest example yet of a price that buys latency rather than capability. Released September 18, 2026, it serves the same weights as GLM-5.3 Flash, the same 320 billion parameter mixture-of-experts activating roughly 18 billion per token, on a faster stack. Z.ai lists it at $0.37 per million input tokens, $0.075 cached, and $1.25 output, which is roughly two and a half times the $0.15 and $0.50 the base Flash model charges, with cache storage free for a limited time. What you get for the premium is throughput: Z.ai advertises up to 200 output tokens per second, and independent measurement on its own API has landed nearer 98. Context is 1 million tokens with up to 131,072 output, and the modality set carries over unchanged, taking text, images, video, and files. The decision rule is simple and unusually clean, because capability is not part of it. If a user is waiting on the stream, FlashX is worth the multiple; if a batch job is running overnight, it is not. Note that FlashX is a hosted tier only and has no separate weight release: the MIT-licensed checkpoint on Hugging Face is GLM-5.3-Flash, and self-hosting it gives you the model but not the serving stack that FlashX is selling.

Input Price

$0.37

per 1M tokens

Output Price

$1.25

per 1M tokens

Context Window

1.0M

tokens

Released

2026-09

API access

Capabilities

textvisionvideotool-usecodereasoning

Key Strengths

  • ✓Up to 200 output tokens per second, vendor-reported
  • ✓Identical 320B MoE weights to GLM-5.3 Flash, so capability is unchanged
  • ✓$0.37/$1.25 with cached input at $0.075
  • ✓1M token context with 131K max output
  • ✓Text, image, video, and file input
  • ✓Cache storage free for a limited time

Best For

  • ▸Interactive coding and chat where a user waits on the stream
  • ▸Latency-sensitive agent loops
  • ▸Computer-use and visual understanding at speed
  • ▸Teams already on GLM-5.3 Flash that need a faster tier without a model change

Pricing Details

Input tokens

$0.37

per 1M tokens

Output tokens

$1.25

per 1M tokens

Estimated cost per 1K requests

$0.99

~1K input + ~500 output tokens avg

Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.

Related Models

View DocumentationCompare ModelsCost CalculatorFull Pricing Guide