Models, measured side by side
Gemma 3 4B vs Qwen3.7 Flash
Compare Gemma 3 4B, Qwen3.7 Flash by token pricing, context length, benchmark results, speed and tool support.
Monthly workload
Monthly cost breakdown
InputOutput
Gemma 3 4B
$1.50 / mo
Qwen3.7 Flash
$1.25 / mo
Performance comparison
Gemma 3 4B
4.8
Qwen3.7 Flash
Not listed
Overview
Pricing & capacity
| Pricing & capacity | Gemma 3 4B | Qwen3.7 Flash |
|---|---|---|
| Input / 1M tokens | $0.05 | $0.03 |
| Output / 1M tokens | $0.10 | $0.13 |
| Cached input / 1M | Not listed | $0.0060 |
| Monthly cost | $1.50 | $1.25 |
| Context window | 131K | 1M |
| Maximum output | 16K | 66K |
Input / 1M tokens
- Gemma 3 4B
- $0.05
- Qwen3.7 Flash
- $0.03
Output / 1M tokens
- Gemma 3 4B
- $0.10
- Qwen3.7 Flash
- $0.13
Cached input / 1M
- Gemma 3 4B
- Not listed
- Qwen3.7 Flash
- $0.0060
Monthly cost
- Gemma 3 4B
- $1.50
- Qwen3.7 Flash
- $1.25
Context window
- Gemma 3 4B
- 131K
- Qwen3.7 Flash
- 1M
Maximum output
- Gemma 3 4B
- 16K
- Qwen3.7 Flash
- 66K
Benchmarks & performance
| Benchmarks & performance | Gemma 3 4B | Qwen3.7 Flash |
|---|---|---|
| Intelligence Index | 4.8 | Not listed |
| GPQA | 29.1% | Not listed |
| Humanity’s Last Exam | 5.3% | Not listed |
Intelligence Index
- Gemma 3 4B
- 4.8
- Qwen3.7 Flash
- Not listed
GPQA
- Gemma 3 4B
- 29.1%
- Qwen3.7 Flash
- Not listed
Humanity’s Last Exam
- Gemma 3 4B
- 5.3%
- Qwen3.7 Flash
- Not listed
Tools & features
| Tools & features | Gemma 3 4B | Qwen3.7 Flash |
|---|---|---|
| Function calling | Not listed | Supported |
| Structured outputs | Supported | Not listed |
| JSON mode | Supported | Supported |
| Reasoning | Not listed | Supported |
| Log probabilities | Not listed | Supported |
| Deterministic seed | Supported | Supported |
| Prompt caching | Not listed | Supported |
Function calling
- Gemma 3 4B
- Not listed
- Qwen3.7 Flash
- Supported
Structured outputs
- Gemma 3 4B
- Supported
- Qwen3.7 Flash
- Not listed
JSON mode
- Gemma 3 4B
- Supported
- Qwen3.7 Flash
- Supported
Reasoning
- Gemma 3 4B
- Not listed
- Qwen3.7 Flash
- Supported
Log probabilities
- Gemma 3 4B
- Not listed
- Qwen3.7 Flash
- Supported
Deterministic seed
- Gemma 3 4B
- Supported
- Qwen3.7 Flash
- Supported
Prompt caching
- Gemma 3 4B
- Not listed
- Qwen3.7 Flash
- Supported
API & availability
| API & availability | Gemma 3 4B | Qwen3.7 Flash |
|---|---|---|
| API identifier | gemma-3-4b-it | qwen3.7-flash |
| Knowledge cutoff | 2024-08-31 | Not listed |
| Prices checked | Oct 8, 2026 | Oct 8, 2026 |
API identifier
- Gemma 3 4B
gemma-3-4b-it- Qwen3.7 Flash
qwen3.7-flash
Knowledge cutoff
- Gemma 3 4B
- 2024-08-31
- Qwen3.7 Flash
- Not listed
Prices checked
- Gemma 3 4B
- Oct 8, 2026
- Qwen3.7 Flash
- Oct 8, 2026
About the models
Gemma 3 4B
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
Full pricing & detailsQwen3.7 Flash
Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...
Full pricing & detailsComparison FAQ
Gemma 3 4B: $0.05 input and $0.10 output per million tokens. Qwen3.7 Flash: $0.03 input and $0.13 output per million tokens. The cheaper choice depends on your input-to-output ratio; a model can have a lower input rate but a higher output rate.
