Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input... Built by Meta.
Prices updated
Input price
$0.10
per 1M tokens · Standard
Output price
$0.30
per 1M tokens · Standard
Input limit
1.3M
tokens
Output limit
16K
tokens
Input formats
Output formats
Llama 4 Scout price history
2 price records since 9/25/2026
Llama 4 Scout cost calculator
$3.50/month
Standard pricing
Overview
What is Llama 4 Scout?
Llama 4 Scout is a text & reasoning and vision model from Meta. Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input... Its 1.3M context window and $0.10 input price make it a candidate for cost-sensitive, high-throughput applications.
Benchmarks
Llama 4 Scout benchmarks & speed
How Llama 4 Scout scores on standardized evaluations, and where it lands among every model we track.
8.1
Intelligence Index
69tok/s
Output speed
793
DesignArena Elo
Reasoning
GPQA Diamond
58.7%
Graduate-level scientific reasoning · top 79%
HLE
3.8%
Humanity's Last Exam · top 92%
Coding
SciCode
21.3%
Python for scientific computing · top 99%
Latency & design
- Time to first token
- 0.86s
- DesignArena win rate
- 26.6%
- Design battles judged
- 1,275
Independent scores from Artificial Analysis and DesignArena · updated 10/8/2026. Higher is better; ranks compare against every model we track with that score.
Rates
Llama 4 Scout pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.10
Per 1M tokens
Cached input · Standard
Not available
No listed cached-input rate
Capabilities
Llama 4 Scout Tools
Tools available when using Llama 4 Scout through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Not supported
Built-in web search
Not supported
Log probabilities
Not supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Not supported
Where to run it
Llama 4 Scout API providers
3 providers serve Llama 4 Scout. Prices are per 1M tokens; uptime is the last 24 hours.
| Provider | Input | Output | Cached | Context | Uptime |
|---|---|---|---|---|---|
| DeepInfrafp8 | $0.10 | $0.30 | — | 328K | 99.97% |
| Novitabf16 | $0.18 | $0.59 | — | 131K | 99.96% |
| $0.25 | $0.70 | — | 1.3M | — |
- Input
- $0.10
- Output
- $0.30
- Cached
- —
- Context
- 328K
- Uptime
- 99.97%
- Input
- $0.18
- Output
- $0.59
- Cached
- —
- Context
- 131K
- Uptime
- 99.96%
- Input
- $0.25
- Output
- $0.70
- Cached
- —
- Context
- 1.3M
- Uptime
- —
Strengths and limitations
Strengths
- Low input cost at $0.10 per million tokens suits high-volume workloads.
- 1.3M context supports large documents, repositories and extended conversations.
- Supports text & reasoning and vision workloads in one model.
Limitations
- Real-world cost still depends on prompt length, response length and provider-specific billing rules.
- Large context capacity does not guarantee consistent retrieval across the entire prompt.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
