gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized... Built by OpenAI.
Prices updated
Input price
$0.15
per 1M tokens · Standard
Output price
$0.60
per 1M tokens · Standard
Input limit
131K
tokens
Output limit
118K
tokens
Input formats
Output formats
gpt-oss-120b price history
2 price records since 9/25/2026
gpt-oss-120b cost calculator
$6.00/month
Standard pricing
Overview
What is gpt-oss-120b?
gpt-oss-120b is a text & reasoning model from OpenAI. gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized... Its 131K context window and $0.15 input price make it a candidate for cost-sensitive, high-throughput applications.
Benchmarks
gpt-oss-120b benchmarks & speed
How gpt-oss-120b scores on standardized evaluations, and where it lands among every model we track.
11.6
Intelligence Index
196tok/s
Output speed
967
DesignArena Elo
Reasoning
GPQA Diamond
78.2%
Graduate-level scientific reasoning · top 57%
HLE
19.6%
Humanity's Last Exam · top 52%
Coding
SciCode
34.0%
Python for scientific computing · top 90%
Latency & design
- Time to first token
- 0.83s
- DesignArena win rate
- 33.4%
- Design battles judged
- 5,272
Independent scores from Artificial Analysis and DesignArena · updated 10/8/2026. Higher is better; ranks compare against every model we track with that score.
Rates
gpt-oss-120b pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.15
Per 1M tokens
Cached input · Standard
$0.07
Per 1M tokens
Capabilities
gpt-oss-120b Tools
Tools available when using gpt-oss-120b through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Supported
Built-in web search
Not supported
Log probabilities
Supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Not supported
Strengths and limitations
Strengths
- Low input cost at $0.15 per million tokens suits high-volume workloads.
- A 131K context window covers most focused application workflows.
- Supports text & reasoning workloads in one model.
Limitations
- Generated tokens cost 4× more than input tokens, which matters for verbose responses.
- The 131K context window is smaller than several long-context alternatives.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
