Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at... Built by Google.
Prices updated
Input price
$0.09
per 1M tokens · Standard
Output price
$0.30
per 1M tokens · Standard
Input limit
262K
tokens
Output limit
236K
tokens
Input formats
Output formats
Gemma 4 26B A4B price history
4 price records since 9/25/2026 · 2 changes
Gemma 4 26B A4B cost calculator
$3.30/month
Standard pricing
Overview
What is Gemma 4 26B A4B?
Gemma 4 26B A4B is a text & reasoning and vision model from Google. Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at... Its 262K context window and $0.09 input price make it a candidate for cost-sensitive, high-throughput applications.
Benchmarks
Gemma 4 26B A4B benchmarks & speed
How Gemma 4 26B A4B scores on standardized evaluations, and where it lands among every model we track.
16.7
Intelligence Index
Reasoning
GPQA Diamond
79.2%
Graduate-level scientific reasoning · top 56%
HLE
19.3%
Humanity's Last Exam · top 52%
Coding
SciCode
40.0%
Python for scientific computing · top 79%
Independent scores from Artificial Analysis and DesignArena · updated 10/5/2026. Higher is better; ranks compare against every model we track with that score.
Rates
Gemma 4 26B A4B pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.09
Per 1M tokens
Cached input · Standard
$0.05
Per 1M tokens
Capabilities
Gemma 4 26B A4B Tools
Tools available when using Gemma 4 26B A4B through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Supported
Built-in web search
Not supported
Log probabilities
Supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Supported
Strengths and limitations
Strengths
- Low input cost at $0.09 per million tokens suits high-volume workloads.
- 262K context supports large documents, repositories and extended conversations.
- Supports text & reasoning and vision workloads in one model.
Limitations
- Generated tokens cost 3× more than input tokens, which matters for verbose responses.
- Large context capacity does not guarantee consistent retrieval across the entire prompt.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
