Models, measured side by side
Gemini 2.5 Pro vs GPT-4.1
Compare Gemini 2.5 Pro, GPT-4.1 by token pricing, context length, benchmark results, speed and tool support.
Monthly workload
Monthly cost breakdown
InputOutput
Gemini 2.5 Pro
$75 / mo
GPT-4.1
$80 / mo
Performance comparison
Gemini 2.5 Pro
16.1
GPT-4.1
12.7
Overview
Pricing & capacity
| Pricing & capacity | Gemini 2.5 Pro | GPT-4.1 |
|---|---|---|
| Input / 1M tokens | $1.25 | $2.00 |
| Output / 1M tokens | $10 | $8.00 |
| Cached input / 1M | $0.13 | $0.50 |
| Monthly cost | $75 | $80 |
| Context window | 1M | 1M |
| Maximum output | 66K | 33K |
Input / 1M tokens
- Gemini 2.5 Pro
- $1.25
- GPT-4.1
- $2.00
Output / 1M tokens
- Gemini 2.5 Pro
- $10
- GPT-4.1
- $8.00
Cached input / 1M
- Gemini 2.5 Pro
- $0.13
- GPT-4.1
- $0.50
Monthly cost
- Gemini 2.5 Pro
- $75
- GPT-4.1
- $80
Context window
- Gemini 2.5 Pro
- 1M
- GPT-4.1
- 1M
Maximum output
- Gemini 2.5 Pro
- 66K
- GPT-4.1
- 33K
Benchmarks & performance
| Benchmarks & performance | Gemini 2.5 Pro | GPT-4.1 |
|---|---|---|
| Arena Elo | 1,255 | 1,200 |
| MMLU-Pro | 82.8% | 73.6% |
| Intelligence Index | 16.1 | 12.7 |
| GPQA | 84.4% | 66.6% |
| Humanity’s Last Exam | 22.5% | 4.2% |
| SciCode | 46.3% | Not listed |
| Output speed | 128 tokens/s | 166 tokens/s |
| Time to first token | 18.08 s | 0.83 s |
| DesignArena rating | 1,156 | 1,029 |
| DesignArena win rate | 57.5% | 50.9% |
Arena Elo
- Gemini 2.5 Pro
- 1,255
- GPT-4.1
- 1,200
MMLU-Pro
- Gemini 2.5 Pro
- 82.8%
- GPT-4.1
- 73.6%
Intelligence Index
- Gemini 2.5 Pro
- 16.1
- GPT-4.1
- 12.7
GPQA
- Gemini 2.5 Pro
- 84.4%
- GPT-4.1
- 66.6%
Humanity’s Last Exam
- Gemini 2.5 Pro
- 22.5%
- GPT-4.1
- 4.2%
SciCode
- Gemini 2.5 Pro
- 46.3%
- GPT-4.1
- Not listed
Output speed
- Gemini 2.5 Pro
- 128 tokens/s
- GPT-4.1
- 166 tokens/s
Time to first token
- Gemini 2.5 Pro
- 18.08 s
- GPT-4.1
- 0.83 s
DesignArena rating
- Gemini 2.5 Pro
- 1,156
- GPT-4.1
- 1,029
DesignArena win rate
- Gemini 2.5 Pro
- 57.5%
- GPT-4.1
- 50.9%
6 of 10
Tools & features
| Tools & features | Gemini 2.5 Pro | GPT-4.1 |
|---|---|---|
| Function calling | Supported | Supported |
| Structured outputs | Supported | Supported |
| JSON mode | Supported | Supported |
| Reasoning | Supported | Not listed |
| Built-in web search | Not listed | Not listed |
| Log probabilities | Not listed | Not listed |
| Deterministic seed | Supported | Supported |
| Parallel tool calls | Not listed | Not listed |
| Prompt caching | Supported | Supported |
Function calling
- Gemini 2.5 Pro
- Supported
- GPT-4.1
- Supported
Structured outputs
- Gemini 2.5 Pro
- Supported
- GPT-4.1
- Supported
JSON mode
- Gemini 2.5 Pro
- Supported
- GPT-4.1
- Supported
Reasoning
- Gemini 2.5 Pro
- Supported
- GPT-4.1
- Not listed
Built-in web search
- Gemini 2.5 Pro
- Not listed
- GPT-4.1
- Not listed
Log probabilities
- Gemini 2.5 Pro
- Not listed
- GPT-4.1
- Not listed
Deterministic seed
- Gemini 2.5 Pro
- Supported
- GPT-4.1
- Supported
Parallel tool calls
- Gemini 2.5 Pro
- Not listed
- GPT-4.1
- Not listed
Prompt caching
- Gemini 2.5 Pro
- Supported
- GPT-4.1
- Supported
6 of 9
API & availability
| API & availability | Gemini 2.5 Pro | GPT-4.1 |
|---|---|---|
| API identifier | gemini-2.5-pro | gpt-4.1 |
| API providers | Google, Google, Google AI Studio, Google, Google AI Studio, Google, Google AI Studio | Azure, OpenAI, Azure |
| Knowledge cutoff | 2025-01-31 | 2024-06-30 |
| Prices checked | Oct 8, 2026 | Oct 8, 2026 |
API identifier
- Gemini 2.5 Pro
gemini-2.5-pro- GPT-4.1
gpt-4.1
API providers
- Gemini 2.5 Pro
- Google, Google, Google AI Studio, Google, Google AI Studio, Google, Google AI Studio
- GPT-4.1
- Azure, OpenAI, Azure
Knowledge cutoff
- Gemini 2.5 Pro
- 2025-01-31
- GPT-4.1
- 2024-06-30
Prices checked
- Gemini 2.5 Pro
- Oct 8, 2026
- GPT-4.1
- Oct 8, 2026
About the models
GPT-4.1
Non-reasoning GPT-4.1 with a 1M-token context window, strong at instruction following and coding.
Full pricing & detailsComparison FAQ
Gemini 2.5 Pro: $1.25 input and $10 output per million tokens. GPT-4.1: $2.00 input and $8.00 output per million tokens. The cheaper choice depends on your input-to-output ratio; a model can have a lower input rate but a higher output rate.
