Google's previous generation Flash model, balancing speed and multimodal capabilities across general agentic and everyday tasks. Built by Google.
Prices updated
Input price
$0.75
per 1M tokens · Standard
Output price
$3.75
per 1M tokens · Standard
Input limit
1M
tokens
Output limit
66K
tokens
Input formats
Output formats
Gemini 3.6 Flash price history
2 price records since 9/25/2026
Gemini 3.6 Flash cost calculator
$34/month
Standard pricing
Overview
What is Gemini 3.6 Flash?
Gemini 3.6 Flash is Google's high-efficiency workhorse model engineered to balance processing speed, reasoning capability, and multimodal ingestion. Positioned within Google's Flash tier, it is optimized to handle multi-step agentic orchestration, code refactoring, and general reasoning at scale. The model natively accepts multimodal inputs across text, code, audio, video, images, and files, backed by an expansive context window of 1,048,576 tokens and a maximum output ceiling of 65,536 tokens. While outputs are limited to text, it integrates developer tooling including structured outputs, response formatting, and implicit context caching to minimize operational latency and token overhead across repetitive API calls.
Benchmarks
Gemini 3.6 Flash benchmarks & speed
How Gemini 3.6 Flash scores on standardized evaluations, and where it lands among every model we track.
34.0
Intelligence Index
192tok/s
Output speed
1293
DesignArena Elo
Reasoning
GPQA Diamond
92.8%
Graduate-level scientific reasoning · top 11%
HLE
40.8%
Humanity's Last Exam · top 20%
Coding
SciCode
53.4%
Python for scientific computing · top 37%
Latency & design
- Time to first token
- 14.58s
- DesignArena win rate
- 53.4%
- Design battles judged
- 30,358
Independent scores from Artificial Analysis and DesignArena · updated 10/8/2026. Higher is better; ranks compare against every model we track with that score.
Rates
Gemini 3.6 Flash pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.75
Per 1M tokens
Cached input · Standard
$0.07
Per 1M tokens
Capabilities
Gemini 3.6 Flash Tools
Tools available when using Gemini 3.6 Flash through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Supported
Built-in web search
Not supported
Log probabilities
Not supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Supported
Where to run it
Gemini 3.6 Flash API providers
7 providers serve Gemini 3.6 Flash. Prices are per 1M tokens; uptime is the last 24 hours.
| Provider | Input | Output | Cached | Context | Uptime |
|---|---|---|---|---|---|
| $0.38 | $1.88 | $0.04 | 1M | 99.34% | |
| Google AI Studio | $0.38 | $1.88 | $0.04 | 1M | 99.85% |
| Google AI Studio | $0.75 | $3.75 | $0.07 | 1M | 99.82% |
| $0.75 | $3.75 | $0.07 | 1M | 99.86% | |
| $1.35 | $6.75 | $0.14 | 1M | 99.83% | |
| Google AI Studio | $1.35 | $6.75 | $0.14 | 1M | 99.68% |
| $0.82 | $4.13 | $0.08 | 1M | 100% |
- Input
- $0.38
- Output
- $1.88
- Cached
- $0.04
- Context
- 1M
- Uptime
- 99.34%
- Input
- $0.38
- Output
- $1.88
- Cached
- $0.04
- Context
- 1M
- Uptime
- 99.85%
- Input
- $0.75
- Output
- $3.75
- Cached
- $0.07
- Context
- 1M
- Uptime
- 99.82%
- Input
- $0.75
- Output
- $3.75
- Cached
- $0.07
- Context
- 1M
- Uptime
- 99.86%
- Input
- $1.35
- Output
- $6.75
- Cached
- $0.14
- Context
- 1M
- Uptime
- 99.83%
- Input
- $1.35
- Output
- $6.75
- Cached
- $0.14
- Context
- 1M
- Uptime
- 99.68%
- Input
- $0.82
- Output
- $4.13
- Cached
- $0.08
- Context
- 1M
- Uptime
- 100%
6 of 7
Strengths and limitations
Strengths
- Native multimodal ingestion across text, audio, video, images, and file uploads.
- Expansive 1,048,576 token context window paired with a large 65,536 token output ceiling.
- Integrated implicit prompt caching that reduces cached input token costs to $0.075 per million tokens.
- Strong support for agentic development with tool calling, structured outputs, and reasoning capabilities.
Limitations
- Produces text outputs only, without native multimodal output generation like image or speech synthesis.
- Priced higher per token than lightweight small-tier alternatives like Flash-Lite.
- Specific training knowledge cutoff information is not publicly documented.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
More models from Google
Agentic Workflows and Multimodal Processing
Gemini 3.6 Flash is engineered for production workloads requiring reliable instruction following and multi-step execution. Featuring support for tool calling, strict response format schema validation, and reasoning capabilities, it handles complex system orchestration and full-stack software refactoring.
Its multimodal input engine accepts arbitrary combinations of video footage, audio streams, high-resolution imagery, and text files. Combined with a 1,048,576 token context window, applications can ingest hours of media or entire codebases in a single API call.
Cost Optimization with Implicit Caching
Gemini 3.6 Flash operates at $0.75 per million input tokens and $3.75 per million output tokens. For high-frequency agent loops and retrieval-augmented generation pipelines that repeatedly submit identical system prompts or reference documents, the model automatically leverages implicit prompt caching.
Cached input tokens are priced at $0.075 per million tokens, offering a 90% discount compared to uncached inputs. This caching mechanism allows developers to maintain rich system context without incurring linear cost growth.
