Grok 4.20
grok-4.20Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering... Built by xAI.
Prices updated
Input price
$1.25
per 1M tokens · Standard
Output price
$2.50
per 1M tokens · Standard
Input limit
2M
tokens
Output limit
1.8M
tokens
Input formats
Output formats
Grok 4.20 price history
2 price records since 9/25/2026
Grok 4.20 cost calculator
$38/month
Standard pricing
Overview
What is Grok 4.20?
Grok 4.20 is a text & reasoning and vision model from xAI. Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering... Its 2M context window and $1.25 input price make it a candidate for quality-focused production applications.
Benchmarks
Grok 4.20 benchmarks & speed
How Grok 4.20 scores on standardized evaluations, and where it lands among every model we track.
25.7
Intelligence Index
120tok/s
Output speed
Reasoning
GPQA Diamond
91.1%
Graduate-level scientific reasoning · top 21%
HLE
34.5%
Humanity's Last Exam · top 32%
Latency & design
- Time to first token
- 19.5s
Independent scores from Artificial Analysis and DesignArena · updated 10/5/2026. Higher is better; ranks compare against every model we track with that score.
Rates
Grok 4.20 pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$1.25
Per 1M tokens
Cached input · Standard
$0.20
Per 1M tokens
Capabilities
Grok 4.20 Tools
Tools available when using Grok 4.20 through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Supported
Built-in web search
Not supported
Log probabilities
Supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Supported
Strengths and limitations
Strengths
- Measured results (AA Intelligence 25.7) position it for demanding production work.
- 2M context supports large documents, repositories and extended conversations.
- Supports text & reasoning and vision workloads in one model.
Limitations
- Real-world cost still depends on prompt length, response length and provider-specific billing rules.
- Large context capacity does not guarantee consistent retrieval across the entire prompt.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
