Compact GPT-5.4 model for fast, affordable everyday tasks. Built by OpenAI.
Prices updated
Input price
$0.75
per 1M tokens · Standard
Output price
$4.50
per 1M tokens · Standard
Input limit
400K
tokens
Output limit
128K
tokens
Input formats
Output formats
GPT-5.4 mini price history
2 price records since 9/25/2026
GPT-5.4 mini cost calculator
$38/month
Standard pricing
Overview
What is GPT-5.4 mini?
GPT-5.4 mini is a text & reasoning and code & slms and vision model from OpenAI. Compact GPT-5.4 model for fast, affordable everyday tasks. Its 400K context window and $0.75 input price make it a candidate for quality-focused production applications.
Benchmarks
GPT-5.4 mini benchmarks & speed
How GPT-5.4 mini scores on standardized evaluations, and where it lands among every model we track.
24.1
Intelligence Index
246tok/s
Output speed
Reasoning
GPQA Diamond
87.5%
Graduate-level scientific reasoning · top 33%
HLE
28.1%
Humanity's Last Exam · top 40%
Coding
SciCode
52.1%
Python for scientific computing · top 39%
Latency & design
- Time to first token
- 129.41s
Independent scores from Artificial Analysis and DesignArena · updated 10/5/2026. Higher is better; ranks compare against every model we track with that score.
Rates
GPT-5.4 mini pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.75
Per 1M tokens
Cached input · Standard
$0.07
Per 1M tokens
Capabilities
GPT-5.4 mini Tools
Tools available when using GPT-5.4 mini through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Supported
Built-in web search
Not supported
Log probabilities
Not supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Supported
Where to run it
GPT-5.4 mini API providers
5 providers serve GPT-5.4 mini. Prices are per 1M tokens; uptime is the last 24 hours.
| Provider | Input | Output | Cached | Context | Uptime |
|---|---|---|---|---|---|
| OpenAI | $0.38 | $2.25 | $0.04 | 400K | 99.98% |
| Azure | $0.75 | $4.50 | $0.07 | 400K | 99.93% |
| OpenAI | $0.75 | $4.50 | $0.07 | 400K | 99.93% |
| Azure | $0.82 | $4.95 | $0.08 | 400K | 100% |
| OpenAI | $1.50 | $9.00 | $0.15 | 400K | 99.87% |
- Input
- $0.38
- Output
- $2.25
- Cached
- $0.04
- Context
- 400K
- Uptime
- 99.98%
- Input
- $0.75
- Output
- $4.50
- Cached
- $0.07
- Context
- 400K
- Uptime
- 99.93%
- Input
- $0.75
- Output
- $4.50
- Cached
- $0.07
- Context
- 400K
- Uptime
- 99.93%
- Input
- $0.82
- Output
- $4.95
- Cached
- $0.08
- Context
- 400K
- Uptime
- 100%
- Input
- $1.50
- Output
- $9.00
- Cached
- $0.15
- Context
- 400K
- Uptime
- 99.87%
Strengths and limitations
Strengths
- Measured results (AA Intelligence 24.1 · Arena Elo 1200 · MMLU-Pro 73.6%) position it for demanding production work.
- 400K context supports large documents, repositories and extended conversations.
- Supports text & reasoning and code & slms and vision workloads in one model.
Limitations
- Generated tokens cost 6× more than input tokens, which matters for verbose responses.
- Large context capacity does not guarantee consistent retrieval across the entire prompt.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
