- Input $/1M
- $1.00
- Output $/1M
- $3.20
- Context
- 200K
- Max output
- 128K
- Reasoning
- Yes
- Included in
- 6 plans
API model id
glm-501 / Standard API priceUSD / 1M tokens
Input
$1.00
Output
$3.20
Cache read
$0.20
Published rate checked 2026-09-29. Sources below.
Current published base rates. Context tiers, cache-write duration and other conditions may change the rate for a request. These are not a reconstructed historical invoice.
02 / Capabilities & limitsProvider specifications
- Context window
- 200,000
- Maximum output
- 128,000
- Input
- text
- Output
- text
- Reasoning
- Supported
- Tool calling
- Supported
- Structured output
- Supported
Thinking is on by default and can be turned off.
Before you chooseChecked 2026-09-29
- Thinking
- Thinking is on by default and can be turned off with thinking.type set to disabled.
- Input support
- Text input only.
- API capabilities
- Function calling, streaming, context caching and structured output are supported.
- Released
- February 12, 2026
- Prompt caching
- Cached input costs $0.20 per MTok. Cached input storage is listed as limited-time free.
- Coding Plan access
- Z.ai's GLM-5 page says GLM-5 is available in the GLM Coding Plan on Pro and Max.
03 / Where you can use it1 API route · 6 subscriptions
Z.AI (Zhipu) APIDirect API
Command Code Go$1 / month ↗Command Code GOAT$10 / month ↗Command Code Pro$20 / month ↗Command Code Max 10×$100 / month ↗Command Code Max 20×$200 / month ↗Kiro Pro$20 / month ↗
Pricing, assumptions & evidence
Published base rate
Input $1.00 · Output $3.20 · Cache read $0.20 · Cache write Not listed
- Official model specifications
- Default thinking behavior
- Official specifications and conditions
- Release date
- Cached input rates
Missing token-category prices are not zero. Workload pricing applies exact recorded categories and admitted routes. No benchmark score or quality ranking is inferred from prices.
Aliases, routes and identity
Catalog ID glm-5
Model release (kind: release)
Developer: Z.AI (Zhipu)
z-ai/glm-5harness_alias

