A cost-efficient model, optimized for high-volume agentic tasks, translation, and simple data processing. Built by Google.
Prices updated
Input price
$0.25
per 1M tokens · Standard
Output price
$1.50
per 1M tokens · Standard
Input limit
1M
tokens
Output limit
66K
tokens
Input formats
Output formats
Gemini 3.1 Flash-Lite price history
2 price records since 9/25/2026
Gemini 3.1 Flash-Lite cost calculator
$13/month
Standard pricing
Overview
What is Gemini 3.1 Flash-Lite?
Gemini 3.1 Flash-Lite is a text & reasoning and vision model from Google. A cost-efficient model, optimized for high-volume agentic tasks, translation, and simple data processing. Its 1M context window and $0.25 input price make it a candidate for cost-sensitive, high-throughput applications.
Rates
Gemini 3.1 Flash-Lite pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.25
Per 1M tokens
Cached input · Standard
$0.03
Per 1M tokens
Capabilities
Gemini 3.1 Flash-Lite Tools
Tools available when using Gemini 3.1 Flash-Lite through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Supported
Built-in web search
Not supported
Log probabilities
Not supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Supported
Where to run it
Gemini 3.1 Flash-Lite API providers
8 providers serve Gemini 3.1 Flash-Lite. Prices are per 1M tokens; uptime is the last 24 hours.
| Provider | Input | Output | Cached | Context | Uptime |
|---|---|---|---|---|---|
| $0.13 | $0.75 | $0.01 | 1M | 100% | |
| Google AI Studio | $0.13 | $0.75 | $0.01 | 1M | 99.98% |
| Google AI Studio | $0.25 | $1.50 | $0.03 | 1M | 99.85% |
| $0.25 | $1.50 | $0.03 | 1M | 99.9% | |
| $0.28 | $1.65 | $0.03 | 1M | 100% | |
| $0.28 | $1.65 | $0.03 | 1M | 99.99% | |
| $0.45 | $2.70 | $0.04 | 1M | 100% | |
| Google AI Studio | $0.45 | $2.70 | $0.04 | 1M | 99.86% |
- Input
- $0.13
- Output
- $0.75
- Cached
- $0.01
- Context
- 1M
- Uptime
- 100%
- Input
- $0.13
- Output
- $0.75
- Cached
- $0.01
- Context
- 1M
- Uptime
- 99.98%
- Input
- $0.25
- Output
- $1.50
- Cached
- $0.03
- Context
- 1M
- Uptime
- 99.85%
- Input
- $0.25
- Output
- $1.50
- Cached
- $0.03
- Context
- 1M
- Uptime
- 99.9%
- Input
- $0.28
- Output
- $1.65
- Cached
- $0.03
- Context
- 1M
- Uptime
- 100%
- Input
- $0.28
- Output
- $1.65
- Cached
- $0.03
- Context
- 1M
- Uptime
- 99.99%
- Input
- $0.45
- Output
- $2.70
- Cached
- $0.04
- Context
- 1M
- Uptime
- 100%
- Input
- $0.45
- Output
- $2.70
- Cached
- $0.04
- Context
- 1M
- Uptime
- 99.86%
6 of 8
Strengths and limitations
Strengths
- Low input cost at $0.25 per million tokens suits high-volume workloads.
- 1M context supports large documents, repositories and extended conversations.
- Supports text & reasoning and vision workloads in one model.
Limitations
- Generated tokens cost 6× more than input tokens, which matters for verbose responses.
- Large context capacity does not guarantee consistent retrieval across the entire prompt.
- Arena Elo and MMLU-Pro are directional; test accuracy, latency and reliability on your own workload before committing.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
