GLM API pricing
Z.ai's GLM models. GLM-5.3 costs $1.40 / $4.40 per 1M tokens; GLM-5.3 Flash costs $0.15 / $0.50.
GLM-5.3
What it costs in practice
GLM-5.3 at list prices, no batch discount.
| 1,000 chat replies1,500 tokens in and 300 out each | $3.42 |
|---|---|
| Summarising 100 long reports20,000 tokens in and 800 out each | $3.15 |
| A month of a coding agent400 tasks, 85% of input cached | $358 |
The coding-agent example assumes 85% of input is served from cache, typical for agents that resend the same context each step. Model your own workload
Details
glm-5.3Checked 27 Sep 2026 against Z.ai pricing and DataNorth (release date) .
Other GLM models
Mistral also resells GLM-5.3 at the same price.
| Model | Input / 1M | Output / 1M | Cached input | Context | Released |
|---|---|---|---|---|---|
| $0.15 | $0.50 | $0.030 | 1M | 26 Aug 2026 |
GLM-5.3 Flash details
GLM-5.3 Flash
$0.15 / $0.50 per 1M tokenscached $0.030 · 1M context
Checked 27 Sep 2026 against Z.ai pricing .
Compare GLM-5.3
Price history
All changesZ.ai pricing rules
- Mistral also resells GLM-5.3 at the same price.
Questions
How much does GLM-5.3 cost?
As of 27 Sep 2026, GLM-5.3 costs $1.40 per million input tokens and $4.40 per million output tokens, and $0.26 per million cached input tokens. Prices checked 27 Sep 2026.
What is the cheapest GLM model?
GLM-5.3 Flash, at $0.15 / $0.50 per million input / output tokens.
What is GLM-5.3's context window?
1M tokens.
Does GLM-5.3 support prompt caching?
Yes. Repeated input read from the cache costs $0.26 per million tokens, 81% less than normal input.