PriceIndex
Google logo

Gemini 3.6 Flash

gemini-3.6-flash

Google's previous generation Flash model, balancing speed and multimodal capabilities across general agentic and everyday tasks. Built by Google.

Prices updated

Input price

$0.75

per 1M tokens · Standard

Output price

$3.75

per 1M tokens · Standard

Input limit

1M

tokens

Output limit

66K

tokens

Input formats

TextImageVideoAudioPDF

Output formats

TextImageVideoAudioPDF
Compare

Gemini 3.6 Flash price history

2 price records since 9/25/2026

Gemini 3.6 Flash cost calculator

Input20M tokens
Output5M tokens

$34/month

Standard pricing

    Overview

    What is Gemini 3.6 Flash?

    Gemini 3.6 Flash is Google's high-efficiency workhorse model engineered to balance processing speed, reasoning capability, and multimodal ingestion. Positioned within Google's Flash tier, it is optimized to handle multi-step agentic orchestration, code refactoring, and general reasoning at scale. The model natively accepts multimodal inputs across text, code, audio, video, images, and files, backed by an expansive context window of 1,048,576 tokens and a maximum output ceiling of 65,536 tokens. While outputs are limited to text, it integrates developer tooling including structured outputs, response formatting, and implicit context caching to minimize operational latency and token overhead across repetitive API calls.

    Multi-step agentic orchestrationFull-stack code refactoringLong-form multimedia document analysisHigh-throughput enterprise RAG

    Benchmarks

    Gemini 3.6 Flash benchmarks & speed

    How Gemini 3.6 Flash scores on standardized evaluations, and where it lands among every model we track.

    34.0

    Intelligence Index

    Better than 79% of 153 models

    192tok/s

    Output speed

    Better than 86% of 119 models

    1293

    DesignArena Elo

    Better than 81% of 81 models

    Reasoning

    GPQA Diamond

    92.8%

    Graduate-level scientific reasoning · top 11%

    HLE

    40.8%

    Humanity's Last Exam · top 20%

    Coding

    SciCode

    53.4%

    Python for scientific computing · top 37%

    Latency & design

    Time to first token
    14.58s
    DesignArena win rate
    53.4%
    Design battles judged
    30,358

    Independent scores from Artificial Analysis and DesignArena · updated 10/8/2026. Higher is better; ranks compare against every model we track with that score.

    Rates

    Gemini 3.6 Flash pricing

    Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.

    Input · Standard

    $0.75

    Per 1M tokens

    Cached input · Standard

    $0.07

    Per 1M tokens

    Capabilities

    Gemini 3.6 Flash Tools

    Tools available when using Gemini 3.6 Flash through supported provider APIs.

    Function calling

    Supported

    Structured outputs

    Supported

    JSON mode

    Supported

    Reasoning

    Supported

    Built-in web search

    Not supported

    Log probabilities

    Not supported

    Deterministic seed

    Supported

    Parallel tool calls

    Not supported

    Prompt caching

    Supported

    Where to run it

    Gemini 3.6 Flash API providers

    7 providers serve Gemini 3.6 Flash. Prices are per 1M tokens; uptime is the last 24 hours.

    Google
    Input
    $0.38
    Output
    $1.88
    Cached
    $0.04
    Context
    1M
    Uptime
    99.34%
    Google AI Studio
    Input
    $0.38
    Output
    $1.88
    Cached
    $0.04
    Context
    1M
    Uptime
    99.85%
    Google AI Studio
    Input
    $0.75
    Output
    $3.75
    Cached
    $0.07
    Context
    1M
    Uptime
    99.82%
    Google
    Input
    $0.75
    Output
    $3.75
    Cached
    $0.07
    Context
    1M
    Uptime
    99.86%
    Google
    Input
    $1.35
    Output
    $6.75
    Cached
    $0.14
    Context
    1M
    Uptime
    99.83%
    Google AI Studio
    Input
    $1.35
    Output
    $6.75
    Cached
    $0.14
    Context
    1M
    Uptime
    99.68%

    6 of 7

    Strengths and limitations

    Strengths

    • Native multimodal ingestion across text, audio, video, images, and file uploads.
    • Expansive 1,048,576 token context window paired with a large 65,536 token output ceiling.
    • Integrated implicit prompt caching that reduces cached input token costs to $0.075 per million tokens.
    • Strong support for agentic development with tool calling, structured outputs, and reasoning capabilities.

    Limitations

    • Produces text outputs only, without native multimodal output generation like image or speech synthesis.
    • Priced higher per token than lightweight small-tier alternatives like Flash-Lite.
    • Specific training knowledge cutoff information is not publicly documented.

    Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.

    Same provider

    More models from Google

    View provider →

    Agentic Workflows and Multimodal Processing

    Gemini 3.6 Flash is engineered for production workloads requiring reliable instruction following and multi-step execution. Featuring support for tool calling, strict response format schema validation, and reasoning capabilities, it handles complex system orchestration and full-stack software refactoring.

    Its multimodal input engine accepts arbitrary combinations of video footage, audio streams, high-resolution imagery, and text files. Combined with a 1,048,576 token context window, applications can ingest hours of media or entire codebases in a single API call.

    Cost Optimization with Implicit Caching

    Gemini 3.6 Flash operates at $0.75 per million input tokens and $3.75 per million output tokens. For high-frequency agent loops and retrieval-augmented generation pipelines that repeatedly submit identical system prompts or reference documents, the model automatically leverages implicit prompt caching.

    Cached input tokens are priced at $0.075 per million tokens, offering a 90% discount compared to uncached inputs. This caching mechanism allows developers to maintain rich system context without incurring linear cost growth.

    Frequently asked questions about Gemini 3.6 Flash

    Gemini 3.6 Flash is a text & reasoning and vision model from Google. Google's previous generation Flash model, balancing speed and multimodal capabilities across general agentic and everyday tasks.