Models, measured side by side
DeepSeek V4.1 Flash vs Qwen3.8 2.4T A95B
Compare DeepSeek V4.1 Flash, Qwen3.8 2.4T A95B by token pricing, context length, benchmark results, speed and tool support.
Monthly workload
Monthly cost breakdown
InputOutput
DeepSeek V4.1 Flash
$7.00 / mo
Qwen3.8 2.4T A95B
$70 / mo
Performance comparison
DeepSeek V4.1 Flash
39.5
Qwen3.8 2.4T A95B
39.9
Overview
Pricing & capacity
| Pricing & capacity | DeepSeek V4.1 Flash | Qwen3.8 2.4T A95B |
|---|---|---|
| Input / 1M tokens | $0.05 | $2.00 |
| Output / 1M tokens | $1.20 | $6.00 |
| Cached input / 1M | $0.02 | $0.25 |
| Monthly cost | $7.00 | $70 |
| Context window | 1M | 1M |
| Maximum output | 944K | 131K |
Input / 1M tokens
- DeepSeek V4.1 Flash
- $0.05
- Qwen3.8 2.4T A95B
- $2.00
Output / 1M tokens
- DeepSeek V4.1 Flash
- $1.20
- Qwen3.8 2.4T A95B
- $6.00
Cached input / 1M
- DeepSeek V4.1 Flash
- $0.02
- Qwen3.8 2.4T A95B
- $0.25
Monthly cost
- DeepSeek V4.1 Flash
- $7.00
- Qwen3.8 2.4T A95B
- $70
Context window
- DeepSeek V4.1 Flash
- 1M
- Qwen3.8 2.4T A95B
- 1M
Maximum output
- DeepSeek V4.1 Flash
- 944K
- Qwen3.8 2.4T A95B
- 131K
Benchmarks & performance
| Benchmarks & performance | DeepSeek V4.1 Flash | Qwen3.8 2.4T A95B |
|---|---|---|
| Intelligence Index | 39.5 | 39.9 |
| GPQA | Not listed | 93.5% |
| Humanity’s Last Exam | 39.2% | 42.4% |
| SciCode | 51.9% | 54.1% |
| Output speed | 213 tokens/s | 38 tokens/s |
| Time to first token | 1.05 s | 2.75 s |
| DesignArena rating | 1,325 | Not listed |
| DesignArena win rate | 52.1% | Not listed |
Intelligence Index
- DeepSeek V4.1 Flash
- 39.5
- Qwen3.8 2.4T A95B
- 39.9
GPQA
- DeepSeek V4.1 Flash
- Not listed
- Qwen3.8 2.4T A95B
- 93.5%
Humanity’s Last Exam
- DeepSeek V4.1 Flash
- 39.2%
- Qwen3.8 2.4T A95B
- 42.4%
SciCode
- DeepSeek V4.1 Flash
- 51.9%
- Qwen3.8 2.4T A95B
- 54.1%
Output speed
- DeepSeek V4.1 Flash
- 213 tokens/s
- Qwen3.8 2.4T A95B
- 38 tokens/s
Time to first token
- DeepSeek V4.1 Flash
- 1.05 s
- Qwen3.8 2.4T A95B
- 2.75 s
DesignArena rating
- DeepSeek V4.1 Flash
- 1,325
- Qwen3.8 2.4T A95B
- Not listed
DesignArena win rate
- DeepSeek V4.1 Flash
- 52.1%
- Qwen3.8 2.4T A95B
- Not listed
Tools & features
| Tools & features | DeepSeek V4.1 Flash | Qwen3.8 2.4T A95B |
|---|---|---|
| Function calling | Supported | Supported |
| Structured outputs | Supported | Supported |
| JSON mode | Supported | Supported |
| Reasoning | Supported | Supported |
| Log probabilities | Supported | Supported |
| Deterministic seed | Supported | Supported |
| Prompt caching | Supported | Supported |
Function calling
- DeepSeek V4.1 Flash
- Supported
- Qwen3.8 2.4T A95B
- Supported
Structured outputs
- DeepSeek V4.1 Flash
- Supported
- Qwen3.8 2.4T A95B
- Supported
JSON mode
- DeepSeek V4.1 Flash
- Supported
- Qwen3.8 2.4T A95B
- Supported
Reasoning
- DeepSeek V4.1 Flash
- Supported
- Qwen3.8 2.4T A95B
- Supported
Log probabilities
- DeepSeek V4.1 Flash
- Supported
- Qwen3.8 2.4T A95B
- Supported
Deterministic seed
- DeepSeek V4.1 Flash
- Supported
- Qwen3.8 2.4T A95B
- Supported
Prompt caching
- DeepSeek V4.1 Flash
- Supported
- Qwen3.8 2.4T A95B
- Supported
API & availability
| API & availability | DeepSeek V4.1 Flash | Qwen3.8 2.4T A95B |
|---|---|---|
| API identifier | deepseek-v4.1-flash | qwen3.8-2.4t-a95b |
| Prices checked | Oct 8, 2026 | Oct 8, 2026 |
API identifier
- DeepSeek V4.1 Flash
deepseek-v4.1-flash- Qwen3.8 2.4T A95B
qwen3.8-2.4t-a95b
Prices checked
- DeepSeek V4.1 Flash
- Oct 8, 2026
- Qwen3.8 2.4T A95B
- Oct 8, 2026
About the models
DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
Full pricing & detailsQwen3.8 2.4T A95B
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...
Full pricing & detailsComparison FAQ
DeepSeek V4.1 Flash: $0.05 input and $1.20 output per million tokens. Qwen3.8 2.4T A95B: $2.00 input and $6.00 output per million tokens. The cheaper choice depends on your input-to-output ratio; a model can have a lower input rate but a higher output rate.
