Low-cost multimodal GPT-4o variant. Built by OpenAI.
Prices updated
Input price
$0.15
per 1M tokens · Standard
Output price
$0.60
per 1M tokens · Standard
Input limit
128K
tokens
Output limit
16K
tokens
Input formats
Output formats
GPT-4o mini price history
12 price records since 4/1/2026 · 10 changes
GPT-4o mini cost calculator
Overview
What is GPT-4o mini?
GPT-4o mini is a text & reasoning and vision model from OpenAI. Low-cost multimodal GPT-4o variant. Its 128K context window and $0.15 input price make it a candidate for cost-sensitive, high-throughput applications.
Benchmarks
GPT-4o mini benchmarks & speed
How GPT-4o mini scores on standardized evaluations, and where it lands among every model we track.
6.7
Intelligence Index
89tok/s
Output speed
Reasoning
GPQA Diamond
42.6%
Graduate-level scientific reasoning · top 92%
HLE
4.2%
Humanity's Last Exam · top 88%
Latency & design
- Time to first token
- 0.88s
Independent scores from Artificial Analysis and DesignArena · updated 10/5/2026. Higher is better; ranks compare against every model we track with that score.
Rates
GPT-4o mini pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.15
Per 1M tokens
Cached input · Standard
$0.07
Per 1M tokens
Capabilities
GPT-4o mini Tools
Tools available when using GPT-4o mini through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Not supported
Built-in web search
Supported
Log probabilities
Supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Supported
Where to run it
GPT-4o mini API providers
3 providers serve GPT-4o mini. Prices are per 1M tokens; uptime is the last 24 hours.
| Provider | Input | Output | Cached | Context | Uptime |
|---|---|---|---|---|---|
| Azure | $0.15 | $0.60 | $0.07 | 128K | 99.93% |
| OpenAI | $0.15 | $0.60 | $0.07 | 128K | 99.86% |
| Azure | $0.17 | $0.66 | $0.08 | 128K | 99.94% |
- Input
- $0.15
- Output
- $0.60
- Cached
- $0.07
- Context
- 128K
- Uptime
- 99.93%
- Input
- $0.15
- Output
- $0.60
- Cached
- $0.07
- Context
- 128K
- Uptime
- 99.86%
- Input
- $0.17
- Output
- $0.66
- Cached
- $0.08
- Context
- 128K
- Uptime
- 99.94%
Strengths and limitations
Strengths
- Low input cost at $0.15 per million tokens suits high-volume workloads.
- A 128K context window covers most focused application workflows.
- Supports text & reasoning and vision workloads in one model.
Limitations
- Generated tokens cost 4× more than input tokens, which matters for verbose responses.
- The 128K context window is smaller than several long-context alternatives.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
More models from OpenAI
Peer set
