Smaller GPT-4.1 with the same 1M context window at a fraction of the cost. Built by OpenAI.
Prices updated
Input price
$0.40
per 1M tokens · Standard
Output price
$1.60
per 1M tokens · Standard
Input limit
1M
tokens
Output limit
33K
tokens
Input formats
Output formats
GPT-4.1 mini price history
2 price records since 9/25/2026
GPT-4.1 mini cost calculator
$16/month
Standard pricing
Overview
What is GPT-4.1 mini?
GPT-4.1 mini is a text & reasoning and code & slms and vision model from OpenAI. Smaller GPT-4.1 with the same 1M context window at a fraction of the cost. Its 1M context window and $0.40 input price make it a candidate for quality-focused production applications.
Benchmarks
GPT-4.1 mini benchmarks & speed
How GPT-4.1 mini scores on standardized evaluations, and where it lands among every model we track.
10.2
Intelligence Index
180tok/s
Output speed
997
DesignArena Elo
Reasoning
GPQA Diamond
66.4%
Graduate-level scientific reasoning · top 75%
HLE
5.0%
Humanity's Last Exam · top 79%
Latency & design
- Time to first token
- 0.69s
- DesignArena win rate
- 47.5%
- Design battles judged
- 1,566
Independent scores from Artificial Analysis and DesignArena · updated 10/8/2026. Higher is better; ranks compare against every model we track with that score.
Rates
GPT-4.1 mini pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.40
Per 1M tokens
Cached input · Standard
$0.10
Per 1M tokens
Capabilities
GPT-4.1 mini Tools
Tools available when using GPT-4.1 mini through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Not supported
Built-in web search
Not supported
Log probabilities
Not supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Supported
Where to run it
GPT-4.1 mini API providers
3 providers serve GPT-4.1 mini. Prices are per 1M tokens; uptime is the last 24 hours.
| Provider | Input | Output | Cached | Context | Uptime |
|---|---|---|---|---|---|
| Azure | $0.40 | $1.60 | $0.10 | 1M | 99.95% |
| OpenAI | $0.40 | $1.60 | $0.10 | 1M | 99.92% |
| Azure | $0.44 | $1.76 | $0.11 | 1M | 100% |
- Input
- $0.40
- Output
- $1.60
- Cached
- $0.10
- Context
- 1M
- Uptime
- 99.95%
- Input
- $0.40
- Output
- $1.60
- Cached
- $0.10
- Context
- 1M
- Uptime
- 99.92%
- Input
- $0.44
- Output
- $1.76
- Cached
- $0.11
- Context
- 1M
- Uptime
- 100%
Strengths and limitations
Strengths
- Measured results (AA Intelligence 10.2 · DesignArena Elo 997 · Arena Elo 1167 · MMLU-Pro 68.1%) position it for demanding production work.
- 1M context supports large documents, repositories and extended conversations.
- Supports text & reasoning and code & slms and vision workloads in one model.
Limitations
- Generated tokens cost 4× more than input tokens, which matters for verbose responses.
- Large context capacity does not guarantee consistent retrieval across the entire prompt.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
