GLM 4.5
glm-4.5GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly... Built by Z.ai.
Prices updated
Input price
$0.60
per 1M tokens · Standard
Output price
$2.20
per 1M tokens · Standard
Input limit
131K
tokens
Output limit
98K
tokens
Input formats
Output formats
GLM 4.5 price history
2 price records since 9/25/2026
GLM 4.5 cost calculator
$23/month
Standard pricing
Overview
What is GLM 4.5?
GLM 4.5 is a text & reasoning model from Z.ai. GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly... Its 131K context window and $0.60 input price make it a candidate for quality-focused production applications.
Benchmarks
GLM 4.5 benchmarks & speed
How GLM 4.5 scores on standardized evaluations, and where it lands among every model we track.
1169
DesignArena Elo
Latency & design
- DesignArena win rate
- 54.4%
- Design battles judged
- 19,653
Independent scores from Artificial Analysis and DesignArena · updated 10/8/2026. Higher is better; ranks compare against every model we track with that score.
Rates
GLM 4.5 pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.60
Per 1M tokens
Cached input · Standard
$0.11
Per 1M tokens
Capabilities
GLM 4.5 Tools
Tools available when using GLM 4.5 through supported provider APIs.
Function calling
Supported
Structured outputs
Not supported
JSON mode
Supported
Reasoning
Supported
Built-in web search
Not supported
Log probabilities
Not supported
Deterministic seed
Not supported
Parallel tool calls
Not supported
Prompt caching
Supported
Strengths and limitations
Strengths
- Measured results (DesignArena Elo 1169) position it for demanding production work.
- A 131K context window covers most focused application workflows.
- Supports text & reasoning workloads in one model.
Limitations
- Generated tokens cost 4× more than input tokens, which matters for verbose responses.
- The 131K context window is smaller than several long-context alternatives.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
