14B small language model, great for edge and cheap tasks. Built by Microsoft.
Prices updated
Input price
$0.07
per 1M tokens · Standard
Output price
$0.14
per 1M tokens · Standard
Input limit
16K
tokens
Output limit
15K
tokens
Input formats
Output formats
Phi-4 price history
12 price records since 4/1/2026 · 10 changes
Overview
What is Phi-4?
Phi-4 is a code & slms and text & reasoning model from Microsoft. 14B small language model, great for edge and cheap tasks. Its 16K context window and $0.07 input price make it a candidate for cost-sensitive, high-throughput applications.
Benchmarks
Phi-4 benchmarks & speed
How Phi-4 scores on standardized evaluations, and where it lands among every model we track.
5.9
Intelligence Index
39tok/s
Output speed
Reasoning
GPQA Diamond
57.5%
Graduate-level scientific reasoning · top 82%
HLE
3.8%
Humanity's Last Exam · top 92%
Latency & design
- Time to first token
- 2.59s
Independent scores from Artificial Analysis and DesignArena · updated 10/5/2026. Higher is better; ranks compare against every model we track with that score.
Rates
Phi-4 pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.07
Per 1M tokens
Cached input · Standard
Not available
No listed cached-input rate
Capabilities
Phi-4 Tools
Tools available when using Phi-4 through supported provider APIs.
Function calling
Not supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Not supported
Built-in web search
Not supported
Log probabilities
Not supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Not supported
Strengths and limitations
Strengths
- Low input cost at $0.07 per million tokens suits high-volume workloads.
- A 16K context window covers most focused application workflows.
- Supports code & slms and text & reasoning workloads in one model.
Limitations
- Real-world cost still depends on prompt length, response length and provider-specific billing rules.
- The 16K context window is smaller than several long-context alternatives.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
More models from Microsoft
Peer set
