PriceIndex
Z.

GLM 4.7 Flash

glm-4.7-flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,... Built by Z.ai.

Prices updated

Input price

$0.06

per 1M tokens · Standard

Output price

$0.40

per 1M tokens · Standard

Input limit

200K

tokens

Output limit

118K

tokens

Input formats

TextImageVideoAudioPDF

Output formats

TextImageVideoAudioPDF
Compare

GLM 4.7 Flash price history

2 price records since 9/25/2026

GLM 4.7 Flash cost calculator

Input20M tokens
Output5M tokens

$3.21/month

Standard pricing

    Overview

    What is GLM 4.7 Flash?

    GLM 4.7 Flash is a text & reasoning model from Z.ai. As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,... Its 200K context window and $0.06 input price make it a candidate for cost-sensitive, high-throughput applications.

    Content workflowsClassification

    Benchmarks

    GLM 4.7 Flash benchmarks & speed

    How GLM 4.7 Flash scores on standardized evaluations, and where it lands among every model we track.

    14.9

    Intelligence Index

    Better than 38% of 153 models

    91tok/s

    Output speed

    Better than 48% of 119 models

    1181

    DesignArena Elo

    Better than 40% of 81 models

    Reasoning

    GPQA Diamond

    58.1%

    Graduate-level scientific reasoning · top 80%

    HLE

    7.6%

    Humanity's Last Exam · top 68%

    Latency & design

    Time to first token
    1.46s
    DesignArena win rate
    53.1%
    Design battles judged
    11,706

    Independent scores from Artificial Analysis and DesignArena · updated 10/8/2026. Higher is better; ranks compare against every model we track with that score.

    Rates

    GLM 4.7 Flash pricing

    Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.

    Input · Standard

    $0.06

    Per 1M tokens

    Cached input · Standard

    Not available

    No listed cached-input rate

    Capabilities

    GLM 4.7 Flash Tools

    Tools available when using GLM 4.7 Flash through supported provider APIs.

    Function calling

    Supported

    Structured outputs

    Supported

    JSON mode

    Supported

    Reasoning

    Supported

    Built-in web search

    Not supported

    Log probabilities

    Supported

    Deterministic seed

    Supported

    Parallel tool calls

    Not supported

    Prompt caching

    Not supported

    Strengths and limitations

    Strengths

    • Low input cost at $0.06 per million tokens suits high-volume workloads.
    • 200K context supports large documents, repositories and extended conversations.
    • Supports text & reasoning workloads in one model.

    Limitations

    • Generated tokens cost 7× more than input tokens, which matters for verbose responses.
    • Large context capacity does not guarantee consistent retrieval across the entire prompt.
    • Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.

    Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.

    Same provider

    More models from Z.ai

    View provider →
    131K

    $0.60 in · $2.20 out

    View model
    131K

    $0.13 in · $0.85 out

    View model
    66K

    $0.60 in · $1.80 out

    View model
    205K

    $0.43 in · $1.75 out

    View model

    Frequently asked questions about GLM 4.7 Flash

    GLM 4.7 Flash is a text & reasoning model from Z.ai. As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...