Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows. Built by Google.
Prices updated
Input price
$0.75
per 1M tokens · Standard
Output price
$3.75
per 1M tokens · Standard
Input limit
1M
tokens
Output limit
66K
tokens
Input formats
Output formats
Gemini 3.8 Flash price history
2 price records since 9/25/2026
Gemini 3.8 Flash cost calculator
$34/month
Standard pricing
Overview
What is Gemini 3.8 Flash?
Gemini 3.8 Flash is Google's multimodal model engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows. Designed as an intelligent workhorse in the Flash line, it balances low latency and cost efficiency with extended reasoning capabilities. The model processes multimodal inputs across text, images, video, audio, and documents, outputting text responses. It provides a context window of up to 1,048,576 tokens and can produce up to 65,536 output tokens. Developer features include tool calling, structured outputs, response format controls, native reasoning, seed parameters, and implicit caching. Pricing is set at $0.75 per million input tokens and $3.75 per million output tokens. Reused contexts that hit Google's implicit caching mechanism are billed at a discounted rate of $0.075 per million tokens.
Benchmarks
Gemini 3.8 Flash benchmarks & speed
How Gemini 3.8 Flash scores on standardized evaluations, and where it lands among every model we track.
40.9
Intelligence Index
239tok/s
Output speed
1309
DesignArena Elo
Reasoning
GPQA Diamond
95.3%
Graduate-level scientific reasoning · top 1%
HLE
47.8%
Humanity's Last Exam · top 7%
Coding
SciCode
56.6%
Python for scientific computing · top 19%
Latency & design
- Time to first token
- 13.45s
- DesignArena win rate
- 50.3%
- Design battles judged
- 43,485
Independent scores from Artificial Analysis and DesignArena · updated 10/8/2026. Higher is better; ranks compare against every model we track with that score.
Rates
Gemini 3.8 Flash pricing
Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.
Input · Standard
$0.75
Per 1M tokens
Cached input · Standard
$0.07
Per 1M tokens
Capabilities
Gemini 3.8 Flash Tools
Tools available when using Gemini 3.8 Flash through supported provider APIs.
Function calling
Supported
Structured outputs
Supported
JSON mode
Supported
Reasoning
Supported
Built-in web search
Not supported
Log probabilities
Not supported
Deterministic seed
Supported
Parallel tool calls
Not supported
Prompt caching
Supported
Where to run it
Gemini 3.8 Flash API providers
6 providers serve Gemini 3.8 Flash. Prices are per 1M tokens; uptime is the last 24 hours.
| Provider | Input | Output | Cached | Context | Uptime |
|---|---|---|---|---|---|
| Google AI Studio | $0.38 | $1.88 | $0.04 | 1M | 99.94% |
| $0.38 | $1.88 | $0.04 | 1M | 98.5% | |
| Google AI Studio | $0.75 | $3.75 | $0.07 | 1M | 99.87% |
| $0.75 | $3.75 | $0.07 | 1M | 98.93% | |
| Google AI Studio | $1.35 | $6.75 | $0.14 | 1M | 99.88% |
| $1.35 | $6.75 | $0.14 | 1M | 99.68% |
- Input
- $0.38
- Output
- $1.88
- Cached
- $0.04
- Context
- 1M
- Uptime
- 99.94%
- Input
- $0.38
- Output
- $1.88
- Cached
- $0.04
- Context
- 1M
- Uptime
- 98.5%
- Input
- $0.75
- Output
- $3.75
- Cached
- $0.07
- Context
- 1M
- Uptime
- 99.87%
- Input
- $0.75
- Output
- $3.75
- Cached
- $0.07
- Context
- 1M
- Uptime
- 98.93%
- Input
- $1.35
- Output
- $6.75
- Cached
- $0.14
- Context
- 1M
- Uptime
- 99.88%
- Input
- $1.35
- Output
- $6.75
- Cached
- $0.14
- Context
- 1M
- Uptime
- 99.68%
Strengths and limitations
Strengths
- Supports an extensive context window of up to 1,048,576 tokens paired with a 65,536 token output ceiling.
- Handles multimodal input ingestion across text, image, video, audio, and file formats.
- Offers low-cost inference at $0.75 per million input tokens and $0.075 per million cached input tokens.
- Features built-in reasoning support, deterministic seed controls, and robust structured outputs.
- Enables efficient agentic execution through implicit caching that slashes repeated context costs.
Limitations
- Produces text outputs only, without native image, video, or audio generation capabilities.
- Carries a higher base token price than earlier lightweight Flash models.
- Intensive reasoning and deep agentic loops can lead to variable completion latency.
Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.
Same provider
More models from Google
Architecture and Multimodal Capabilities
Gemini 3.8 Flash handles diverse multimodal inputs spanning text, images, video, audio, and documents. With an input capacity of 1,048,576 tokens and output support up to 65,536 tokens, it allows developers to ingest large software repositories, extensive media streams, and long technical documents in a single query.
In addition to its high capacity, the model incorporates developer-focused controls such as tool use, structured outputs, response format enforcement, and seed parameters. These capabilities, coupled with native reasoning, make it well-suited for autonomous agentic environments requiring systematic, multi-step execution.
Pricing and Cache Efficiency
Google rates Gemini 3.8 Flash at $0.75 per million input tokens and $3.75 per million output tokens. This pricing tier positions the model as an economical option for high-volume production systems that require stronger reasoning than earlier Flash versions.
Operational expenses can be significantly reduced using implicit prompt caching. Static system instructions, tool definitions, and reference documentation stored in cache are billed at $0.075 per million tokens, representing a 90 percent reduction in input token costs for recurring workflows.
