PriceIndex
Google logo

Gemini 3.7 Flash

gemini-3.7-flash

Google's high-speed, efficient Flash model built for everyday coding, agentic tool use, and reliable multi-step execution. Built by Google.

Prices updated

Input price

$0.75

per 1M tokens · Standard

Output price

$3.75

per 1M tokens · Standard

Input limit

1M

tokens

Output limit

66K

tokens

Input formats

TextImageVideoAudioPDF

Output formats

TextImageVideoAudioPDF
Compare

Gemini 3.7 Flash price history

2 price records since 9/25/2026

Gemini 3.7 Flash cost calculator

Input20M tokens
Output5M tokens

$34/month

Standard pricing

    Overview

    What is Gemini 3.7 Flash?

    Gemini 3.7 Flash is Google's multimodal model engineered for high-throughput operational tasks, software development, and multi-step agentic execution. It supports extensive input modalities, including text, audio, video, images, and document files, returning text outputs alongside built-in reasoning capabilities. With a 1,048,576 token context window and a maximum output ceiling of 65,536 tokens, the model is architected for large-scale document comprehension, code refactoring, and agent orchestration. It features native support for tool use, structured JSON outputs, deterministic sampling via seed controls, and implicit caching. Pricing is structured at $0.75 per million input tokens and $3.75 per million output tokens, with cached input queries discounted to $0.075 per million tokens. This low-latency efficiency profile makes it a viable workhorse model for production pipelines requiring fast multi-turn interactions.

    Multi-step agent workflowsFull-stack code refactoringLong-document analysisMultimodal data extractionHigh-volume customer support

    Benchmarks

    Gemini 3.7 Flash benchmarks & speed

    How Gemini 3.7 Flash scores on standardized evaluations, and where it lands among every model we track.

    39.1

    Intelligence Index

    Better than 86% of 153 models

    280tok/s

    Output speed

    Better than 98% of 119 models

    1313

    DesignArena Elo

    Better than 88% of 81 models

    Reasoning

    GPQA Diamond

    94.5%

    Graduate-level scientific reasoning · top 2%

    HLE

    47.9%

    Humanity's Last Exam · top 7%

    Coding

    SciCode

    57.2%

    Python for scientific computing · top 17%

    Latency & design

    Time to first token
    10.62s
    DesignArena win rate
    54%
    Design battles judged
    74,822

    Independent scores from Artificial Analysis and DesignArena · updated 10/8/2026. Higher is better; ranks compare against every model we track with that score.

    Rates

    Gemini 3.7 Flash pricing

    Pricing is based on token usage. Provider-specific caching, batch, regional, and tool charges may affect the final cost.

    Input · Standard

    $0.75

    Per 1M tokens

    Cached input · Standard

    $0.07

    Per 1M tokens

    Capabilities

    Gemini 3.7 Flash Tools

    Tools available when using Gemini 3.7 Flash through supported provider APIs.

    Function calling

    Supported

    Structured outputs

    Supported

    JSON mode

    Supported

    Reasoning

    Supported

    Built-in web search

    Not supported

    Log probabilities

    Not supported

    Deterministic seed

    Supported

    Parallel tool calls

    Not supported

    Prompt caching

    Supported

    Where to run it

    Gemini 3.7 Flash API providers

    6 providers serve Gemini 3.7 Flash. Prices are per 1M tokens; uptime is the last 24 hours.

    Google
    Input
    $0.38
    Output
    $1.88
    Cached
    $0.04
    Context
    1M
    Uptime
    97.59%
    Google AI Studio
    Input
    $0.38
    Output
    $1.88
    Cached
    $0.04
    Context
    1M
    Uptime
    99.94%
    Google AI Studio
    Input
    $0.75
    Output
    $3.75
    Cached
    $0.07
    Context
    1M
    Uptime
    99.77%
    Google
    Input
    $0.75
    Output
    $3.75
    Cached
    $0.07
    Context
    1M
    Uptime
    98.94%
    Google
    Input
    $1.35
    Output
    $6.75
    Cached
    $0.14
    Context
    1M
    Uptime
    99.87%
    Google AI Studio
    Input
    $1.35
    Output
    $6.75
    Cached
    $0.14
    Context
    1M
    Uptime
    99.88%

    Strengths and limitations

    Strengths

    • Massive 1,048,576 token context window supports repository-wide code analysis and comprehensive multi-document processing.
    • Generates up to 65,536 tokens per request, enabling complete code generation and extensive long-form synthesis.
    • Native multimodal ingestion handles text, audio, video, images, and file inputs within a single prompt pipeline.
    • Implicit prompt caching reduces recurring input token costs by up to 90% at $0.075 per million cached tokens.
    • Built-in reasoning and tool-calling capabilities facilitate reliable multi-step agent workflows and structured responses.

    Limitations

    • Output generation is strictly limited to text formats, lacking native generation of image, video, or audio media.
    • Fine-tuning capabilities and custom model training support are not publicly documented.
    • Complex reasoning execution can introduce added latency compared to lightweight, non-reasoning models.

    Arena Elo, MMLU-Pro and pricing figures are illustrative; validate current vendor terms before purchase.

    Same provider

    More models from Google

    View provider →

    Agentic Tool Use and Reasoning Architecture

    Gemini 3.7 Flash is engineered specifically for autonomous developer agents and operational orchestration. By combining built-in reasoning mechanisms with function calling and structured outputs, the model evaluates intermediate states before executing external tools. This deliberate decision-making process reduces tool hallucination and enhances accuracy when navigating APIs, command-line interfaces, and multi-step business logic.

    The model's 65,536 maximum output token capacity also allows it to produce exhaustive, multi-file code modifications and detailed multi-turn execution traces without hitting generation cutoffs mid-task.

    Context Efficiency and Implicit Caching Economics

    Operating across a 1,048,576 token context window allows developers to ingest multimodal inputs such as large video feeds, audio recordings, and complex enterprise documentation in a single session. Standard input tokens are priced at $0.75 per million tokens, while output generation costs $3.75 per million tokens.

    To keep high-volume production systems economical, Gemini 3.7 Flash incorporates implicit caching. When repetitive contexts—such as system prompts, API definitions, or large reference files—are reused across calls, input pricing drops by 90% to $0.075 per million tokens, significantly lowering sustained deployment overhead.

    Frequently asked questions about Gemini 3.7 Flash

    Gemini 3.7 Flash is a text & reasoning and vision and code & slms model from Google. Google's high-speed, efficient Flash model built for everyday coding, agentic tool use, and reliable multi-step execution.