- Input $/1M
- $1.40
- Output $/1M
- $4.40
- Context
- 1M
- Max output
- 128K
- Reasoning
- Yes
- Included in
- 9 plans
API model id
glm-5.201 / Standard API priceUSD / 1M tokens
Input
$1.40
Output
$4.40
Cache read
$0.26
Published rate checked 2026-09-29. Sources below.
Current published base rates. Context tiers, cache-write duration and other conditions may change the rate for a request. These are not a reconstructed historical invoice.
02 / Capabilities & limitsProvider specifications
- Context window
- 1,000,000
- Maximum output
- 128,000
- Input
- text
- Output
- text
- Reasoning
- Supported
- Tool calling
- Supported
- Structured output
- Supported
Thinking is on by default and can be turned off.
Before you chooseChecked 2026-09-29
- Thinking
- Thinking is on by default and can be turned off with thinking.type set to disabled.
- Input support
- Text input only.
- API capabilities
- Function calling, streaming, context caching and structured output are supported.
- Released
- June 16, 2026
- Prompt caching
- Cached input costs $0.26 per MTok. Cached input storage is listed as limited-time free.
03 / Where you can use it1 API route · 9 subscriptions
Z.AI (Zhipu) APIDirect API
Command Code Go$1 / month ↗Command Code GOAT$10 / month ↗Command Code Pro$20 / month ↗Command Code Max 10×$100 / month ↗Command Code Max 20×$200 / month ↗Ollama Cloud Pro$20 / month ↗Ollama Cloud Max$100 / month ↗OpenCode Go$10 / month ↗OpenCode Go Plus$40 / month ↗
Pricing, assumptions & evidence
Published base rate
Input $1.40 · Output $4.40 · Cache read $0.26 · Cache write Not listed
- Official model specifications
- Default thinking behavior
- Official specifications and conditions
- Release date
- Cached input rates
Missing token-category prices are not zero. Workload pricing applies exact recorded categories and admitted routes. No benchmark score or quality ranking is inferred from prices.
Aliases, routes and identity
Catalog ID glm-5-2
Model release (kind: release)
Developer: Z.AI (Zhipu)
z-ai/glm-5.2harness_alias

