GLM 4.7 Flash
glm-4.7-flashAs a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,... Built by Z.ai.
Prices updated
Input price
$0.06
per 1M tokens · Standard
Output price
$0.40
per 1M tokens · Standard
Input limit
200K
tokens
Output limit
118K
tokens
Input formats
Output formats
GLM 4.7 Flash price history
2 price records since 9/25/2026
GLM 4.7 Flash cost calculator
$3.21/month
Standard pricing
Overview
What is GLM 4.7 Flash?
GLM 4.7 Flash is a text & reasoning model from Z.ai. As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,... Its 200K context window and $0.06 input price make it a candidate for cost-sensitive, high-throughput applications.
Benchmarks
GLM 4.7 Flash benchmarks & speed
How GLM 4.7 Flash scores on standardized evaluations, and where it lands among every model we track.
14.9
Intelligence Index
91tok/s
Output speed
1181
DesignArena Elo
Reasoning
GPQA Diamond
58.1%
Graduate-level scientific reasoning · top 80%
HLE
7.6%
Humanity's Last Exam · top 68%
Latency & design
- Time to first token
- 1.46s
- DesignArena win rate
- 53.1%
- Design battles judged
- 11,706
Independent scores from Artificial Analysis and DesignArena · updated 10/8/2026. Higher is better; ranks compare against every model we track with that score.
Rates
GLM 4.7 Flash pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.06
Per 1M tokens
Cached input · Standard
Not available
No listed cached-input rate
Capabilities
GLM 4.7 Flash Tools
Tools available when using GLM 4.7 Flash through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Supported
Built-in web search
Not supported
Log probabilities
Supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Not supported
Strengths and limitations
Strengths
- Low input cost at $0.06 per million tokens suits high-volume workloads.
- 200K context supports large documents, repositories and extended conversations.
- Supports text & reasoning workloads in one model.
Limitations
- Generated tokens cost 7× more than input tokens, which matters for verbose responses.
- Large context capacity does not guarantee consistent retrieval across the entire prompt.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
