Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward... Built by Meta.
Prices updated
Input price
$0.19
per 1M tokens · Standard
Output price
$0.65
per 1M tokens · Standard
Input limit
1M
tokens
Output limit
16K
tokens
Input formats
Output formats
Llama 4 Maverick price history
2 price records since 9/25/2026
Llama 4 Maverick cost calculator
$7.01/month
Standard pricing
Overview
What is Llama 4 Maverick?
Llama 4 Maverick is a text & reasoning and vision model from Meta. Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward... Its 1M context window and $0.19 input price make it a candidate for cost-sensitive, high-throughput applications.
Benchmarks
Llama 4 Maverick benchmarks & speed
How Llama 4 Maverick scores on standardized evaluations, and where it lands among every model we track.
10.0
Intelligence Index
68tok/s
Output speed
883
DesignArena Elo
Reasoning
GPQA Diamond
67.1%
Graduate-level scientific reasoning · top 73%
HLE
4.9%
Humanity's Last Exam · top 81%
Coding
SciCode
31.7%
Python for scientific computing · top 94%
Latency & design
- Time to first token
- 1.24s
- DesignArena win rate
- 35.8%
- Design battles judged
- 1,678
Independent scores from Artificial Analysis and DesignArena · updated 10/8/2026. Higher is better; ranks compare against every model we track with that score.
Rates
Llama 4 Maverick pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.19
Per 1M tokens
Cached input · Standard
$0.05
Per 1M tokens
Capabilities
Llama 4 Maverick Tools
Tools available when using Llama 4 Maverick through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Not supported
Built-in web search
Not supported
Log probabilities
Supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Supported
Strengths and limitations
Strengths
- Low input cost at $0.19 per million tokens suits high-volume workloads.
- 1M context supports large documents, repositories and extended conversations.
- Supports text & reasoning and vision workloads in one model.
Limitations
- Generated tokens cost 3× more than input tokens, which matters for verbose responses.
- Large context capacity does not guarantee consistent retrieval across the entire prompt.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
