Google's high-speed, efficient Flash model built for everyday coding, agentic tool use, and reliable multi-step execution. Built by Google.
Prices updated
Input price
$0.75
per 1M tokens · Standard
Output price
$3.75
per 1M tokens · Standard
Input limit
1M
tokens
Output limit
66K
tokens
Input formats
Output formats
Gemini 3.7 Flash price history
2 price records since 9/25/2026
Gemini 3.7 Flash cost calculator
$34/month
Standard pricing
Overview
What is Gemini 3.7 Flash?
Gemini 3.7 Flash is Google's multimodal model engineered for high-throughput operational tasks, software development, and multi-step agentic execution. It supports extensive input modalities, including text, audio, video, images, and document files, returning text outputs alongside built-in reasoning capabilities. With a 1,048,576 token context window and a maximum output ceiling of 65,536 tokens, the model is architected for large-scale document comprehension, code refactoring, and agent orchestration. It features native support for tool use, structured JSON outputs, deterministic sampling via seed controls, and implicit caching. Pricing is structured at $0.75 per million input tokens and $3.75 per million output tokens, with cached input queries discounted to $0.075 per million tokens. This low-latency efficiency profile makes it a viable workhorse model for production pipelines requiring fast multi-turn interactions.
Benchmarks
Gemini 3.7 Flash benchmarks & speed
How Gemini 3.7 Flash scores on standardized evaluations, and where it lands among every model we track.
39.1
Intelligence Index
280tok/s
Output speed
1313
DesignArena Elo
Reasoning
GPQA Diamond
94.5%
Graduate-level scientific reasoning · top 2%
HLE
47.9%
Humanity's Last Exam · top 7%
Coding
SciCode
57.2%
Python for scientific computing · top 17%
Latency & design
- Time to first token
- 10.62s
- DesignArena win rate
- 54%
- Design battles judged
- 74,822
Independent scores from Artificial Analysis and DesignArena · updated 10/8/2026. Higher is better; ranks compare against every model we track with that score.
Rates
Gemini 3.7 Flash pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.75
Per 1M tokens
Cached input · Standard
$0.07
Per 1M tokens
Capabilities
Gemini 3.7 Flash Tools
Tools available when using Gemini 3.7 Flash through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Supported
Built-in web search
Not supported
Log probabilities
Not supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Supported
Where to run it
Gemini 3.7 Flash API providers
6 providers serve Gemini 3.7 Flash. Prices are per 1M tokens; uptime is the last 24 hours.
| Provider | Input | Output | Cached | Context | Uptime |
|---|---|---|---|---|---|
| $0.38 | $1.88 | $0.04 | 1M | 97.59% | |
| Google AI Studio | $0.38 | $1.88 | $0.04 | 1M | 99.94% |
| Google AI Studio | $0.75 | $3.75 | $0.07 | 1M | 99.77% |
| $0.75 | $3.75 | $0.07 | 1M | 98.94% | |
| $1.35 | $6.75 | $0.14 | 1M | 99.87% | |
| Google AI Studio | $1.35 | $6.75 | $0.14 | 1M | 99.88% |
- Input
- $0.38
- Output
- $1.88
- Cached
- $0.04
- Context
- 1M
- Uptime
- 97.59%
- Input
- $0.38
- Output
- $1.88
- Cached
- $0.04
- Context
- 1M
- Uptime
- 99.94%
- Input
- $0.75
- Output
- $3.75
- Cached
- $0.07
- Context
- 1M
- Uptime
- 99.77%
- Input
- $0.75
- Output
- $3.75
- Cached
- $0.07
- Context
- 1M
- Uptime
- 98.94%
- Input
- $1.35
- Output
- $6.75
- Cached
- $0.14
- Context
- 1M
- Uptime
- 99.87%
- Input
- $1.35
- Output
- $6.75
- Cached
- $0.14
- Context
- 1M
- Uptime
- 99.88%
Strengths and limitations
Strengths
- Massive 1,048,576 token context window supports repository-wide code analysis and comprehensive multi-document processing.
- Generates up to 65,536 tokens per request, enabling complete code generation and extensive long-form synthesis.
- Native multimodal ingestion handles text, audio, video, images, and file inputs within a single prompt pipeline.
- Implicit prompt caching reduces recurring input token costs by up to 90% at $0.075 per million cached tokens.
- Built-in reasoning and tool-calling capabilities facilitate reliable multi-step agent workflows and structured responses.
Limitations
- Output generation is strictly limited to text formats, lacking native generation of image, video, or audio media.
- Fine-tuning capabilities and custom model training support are not publicly documented.
- Complex reasoning execution can introduce added latency compared to lightweight, non-reasoning models.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
More models from Google
Agentic Tool Use and Reasoning Architecture
Gemini 3.7 Flash is engineered specifically for autonomous developer agents and operational orchestration. By combining built-in reasoning mechanisms with function calling and structured outputs, the model evaluates intermediate states before executing external tools. This deliberate decision-making process reduces tool hallucination and enhances accuracy when navigating APIs, command-line interfaces, and multi-step business logic.
The model's 65,536 maximum output token capacity also allows it to produce exhaustive, multi-file code modifications and detailed multi-turn execution traces without hitting generation cutoffs mid-task.
Context Efficiency and Implicit Caching Economics
Operating across a 1,048,576 token context window allows developers to ingest multimodal inputs such as large video feeds, audio recordings, and complex enterprise documentation in a single session. Standard input tokens are priced at $0.75 per million tokens, while output generation costs $3.75 per million tokens.
To keep high-volume production systems economical, Gemini 3.7 Flash incorporates implicit caching. When repetitive contexts—such as system prompts, API definitions, or large reference files—are reused across calls, input pricing drops by 90% to $0.075 per million tokens, significantly lowering sustained deployment overhead.
