Nemotron 3 Ultra
nemotron-3-ultra-550b-a55bNVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it... Built by NVIDIA.
Prices updated
Input price
$0.50
per 1M tokens · Standard
Output price
$2.20
per 1M tokens · Standard
Input limit
262K
tokens
Output limit
16K
tokens
Input formats
Output formats
Nemotron 3 Ultra price history
4 price records since 9/25/2026 · 2 changes
Nemotron 3 Ultra cost calculator
$21/month
Standard pricing
Overview
What is Nemotron 3 Ultra?
Nemotron 3 Ultra is a text & reasoning model from NVIDIA. NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it... Its 262K context window and $0.50 input price make it a candidate for quality-focused production applications.
Benchmarks
Nemotron 3 Ultra benchmarks & speed
How Nemotron 3 Ultra scores on standardized evaluations, and where it lands among every model we track.
1148
DesignArena Elo
Latency & design
- DesignArena win rate
- 36.2%
- Design battles judged
- 23,996
Independent scores from Artificial Analysis and DesignArena · updated 10/8/2026. Higher is better; ranks compare against every model we track with that score.
Rates
Nemotron 3 Ultra pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.50
Per 1M tokens
Cached input · Standard
$0.10
Per 1M tokens
Capabilities
Nemotron 3 Ultra Tools
Tools available when using Nemotron 3 Ultra through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Supported
Built-in web search
Not supported
Log probabilities
Not supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Supported
Where to run it
Nemotron 3 Ultra API providers
4 providers serve Nemotron 3 Ultra. Prices are per 1M tokens; uptime is the last 24 hours.
| Provider | Input | Output | Cached | Context | Uptime |
|---|---|---|---|---|---|
| DeepInfrafp4 | $0.50 | $2.20 | $0.10 | 262K | 99.95% |
| BaseTenfp4 | $0.60 | $2.40 | $0.12 | 203K | 99.99% |
| BaseTenfp4 | $0.60 | $2.40 | $0.12 | 203K | 99.98% |
| Venicefp8 | $0.63 | $3.13 | $0.19 | 256K | 96.9% |
- Input
- $0.50
- Output
- $2.20
- Cached
- $0.10
- Context
- 262K
- Uptime
- 99.95%
- Input
- $0.60
- Output
- $2.40
- Cached
- $0.12
- Context
- 203K
- Uptime
- 99.99%
- Input
- $0.60
- Output
- $2.40
- Cached
- $0.12
- Context
- 203K
- Uptime
- 99.98%
- Input
- $0.63
- Output
- $3.13
- Cached
- $0.19
- Context
- 256K
- Uptime
- 96.9%
Strengths and limitations
Strengths
- Measured results (DesignArena Elo 1148) position it for demanding production work.
- 262K context supports large documents, repositories and extended conversations.
- Supports text & reasoning workloads in one model.
Limitations
- Generated tokens cost 4× more than input tokens, which matters for verbose responses.
- Large context capacity does not guarantee consistent retrieval across the entire prompt.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
