PriceIndex
NV

Nemotron 3 Ultra

nemotron-3-ultra-550b-a55b

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it... Built by NVIDIA.

Prices updated

Input price

$0.50

per 1M tokens · Standard

Output price

$2.20

per 1M tokens · Standard

Input limit

262K

tokens

Output limit

16K

tokens

Input formats

TextImageVideoAudioPDF

Output formats

TextImageVideoAudioPDF
Compare

Nemotron 3 Ultra price history

4 price records since 9/25/2026 · 2 changes

Nemotron 3 Ultra cost calculator

Input20M tokens
Output5M tokens

$21/month

Standard pricing

    Overview

    What is Nemotron 3 Ultra?

    Nemotron 3 Ultra is a text & reasoning model from NVIDIA. NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it... Its 262K context window and $0.50 input price make it a candidate for quality-focused production applications.

    Content workflowsClassification

    Benchmarks

    Nemotron 3 Ultra benchmarks & speed

    How Nemotron 3 Ultra scores on standardized evaluations, and where it lands among every model we track.

    1148

    DesignArena Elo

    Better than 28% of 81 models

    Latency & design

    DesignArena win rate
    36.2%
    Design battles judged
    23,996

    Independent scores from Artificial Analysis and DesignArena · updated 10/8/2026. Higher is better; ranks compare against every model we track with that score.

    Rates

    Nemotron 3 Ultra pricing

    Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.

    Input · Standard

    $0.50

    Per 1M tokens

    Cached input · Standard

    $0.10

    Per 1M tokens

    Capabilities

    Nemotron 3 Ultra Tools

    Tools available when using Nemotron 3 Ultra through supported provider APIs.

    Function calling

    Supported

    Structured outputs

    Supported

    JSON mode

    Supported

    Reasoning

    Supported

    Built-in web search

    Not supported

    Log probabilities

    Not supported

    Deterministic seed

    Supported

    Parallel tool calls

    Not supported

    Prompt caching

    Supported

    Where to run it

    Nemotron 3 Ultra API providers

    4 providers serve Nemotron 3 Ultra. Prices are per 1M tokens; uptime is the last 24 hours.

    DeepInfrafp4
    Input
    $0.50
    Output
    $2.20
    Cached
    $0.10
    Context
    262K
    Uptime
    99.95%
    BaseTenfp4
    Input
    $0.60
    Output
    $2.40
    Cached
    $0.12
    Context
    203K
    Uptime
    99.99%
    BaseTenfp4
    Input
    $0.60
    Output
    $2.40
    Cached
    $0.12
    Context
    203K
    Uptime
    99.98%
    Venicefp8
    Input
    $0.63
    Output
    $3.13
    Cached
    $0.19
    Context
    256K
    Uptime
    96.9%

    Strengths and limitations

    Strengths

    • Measured results (DesignArena Elo 1148) position it for demanding production work.
    • 262K context supports large documents, repositories and extended conversations.
    • Supports text & reasoning workloads in one model.

    Limitations

    • Generated tokens cost 4× more than input tokens, which matters for verbose responses.
    • Large context capacity does not guarantee consistent retrieval across the entire prompt.
    • Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.

    Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.

    Same provider

    More models from NVIDIA

    View provider →

    $0.20 in · $0.20 out

    View model

    $0.05 in · $0.14 out

    View model

    $0.06 in · $0.24 out

    View model

    $0.09 in · $0.40 out

    View model

    Frequently asked questions about Nemotron 3 Ultra

    Nemotron 3 Ultra is a text & reasoning model from NVIDIA. NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...