DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context... Built by DeepSeek.
Prices updated
Input price
$0.25
per 1M tokens · Standard
Output price
$0.95
per 1M tokens · Standard
Input limit
164K
tokens
Output limit
33K
tokens
Input formats
Output formats
DeepSeek V3.1 price history
2 price records since 9/25/2026
DeepSeek V3.1 cost calculator
$9.75/month
Standard pricing
Overview
What is DeepSeek V3.1?
DeepSeek V3.1 is a text & reasoning model from DeepSeek. DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context... Its 164K context window and $0.25 input price make it a candidate for cost-sensitive, high-throughput applications.
Benchmarks
DeepSeek V3.1 benchmarks & speed
How DeepSeek V3.1 scores on standardized evaluations, and where it lands among every model we track.
13.7
Intelligence Index
Reasoning
GPQA Diamond
73.5%
Graduate-level scientific reasoning · top 67%
HLE
6.7%
Humanity's Last Exam · top 72%
Independent scores from Artificial Analysis and DesignArena · updated 10/5/2026. Higher is better; ranks compare against every model we track with that score.
Rates
DeepSeek V3.1 pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.25
Per 1M tokens
Cached input · Standard
$0.13
Per 1M tokens
Capabilities
DeepSeek V3.1 Tools
Tools available when using DeepSeek V3.1 through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Supported
Built-in web search
Not supported
Log probabilities
Supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Supported
Strengths and limitations
Strengths
- Low input cost at $0.25 per million tokens suits high-volume workloads.
- A 164K context window covers most focused application workflows.
- Supports text & reasoning workloads in one model.
Limitations
- Generated tokens cost 4× more than input tokens, which matters for verbose responses.
- The 164K context window is smaller than several long-context alternatives.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
