PriceIndex

Models, measured side by side

Gemini 3.8 Flash vs GPT-6 Luna

Compare Gemini 3.8 Flash, GPT-6 Luna by token pricing, context length, benchmark results, speed and tool support.

Updated Oct 2026 · Last verified Oct 9, 2026 · Public listed prices. Confirm with the provider.

Verdict

Gemini 3.8 Flash vs GPT-6 Luna
Cheaper

GPT-6 Luna

87% less than next at 3:1

Smarter

Gemini 3.8 Flash

AA 40.9 vs 38.1

Bigger context

GPT-6 Luna

1.1M vs 1M

Choose Gemini 3.8 Flash if…

  • you need the strongest reasoning (AA Intelligence 40.9 vs 38.1)

Choose GPT-6 Luna if…

  • you want the lowest cost on input-heavy work (87% cheaper at 3:1)
  • you generate a lot of output ($0.500 vs $3.75 per 1M output tokens)
  • you need the biggest context window (1.1M vs 1M tokens)
10M input + 2M output / monthGemini 3.8 Flash: $15.00GPT-6 Luna: $2.00

Monthly workload

Monthly cost breakdown

InputOutput

Gemini 3.8 Flash

$34 / mo

GPT-6 Luna

$4.50 / mo

Performance comparison

Gemini 3.8 Flash

40.9

GPT-6 Luna

38.1

Overview

Released
Gemini 3.8 Flash
Not listed
GPT-6 Luna
2026-08
Status
Gemini 3.8 Flash
Active
GPT-6 Luna
Active
Input formats
Gemini 3.8 Flash
textimagevideofileaudio
GPT-6 Luna
fileimagetext
Output formats
Gemini 3.8 Flash
text
GPT-6 Luna
text

Pricing & capacity

Input / 1M tokens
Gemini 3.8 Flash
$0.75
GPT-6 Luna
$0.10
Output / 1M tokens
Gemini 3.8 Flash
$3.75
GPT-6 Luna
$0.50
Cached input / 1M
Gemini 3.8 Flash
$0.07
GPT-6 Luna
$0.01
Monthly cost
Gemini 3.8 Flash
$34
GPT-6 Luna
$4.50
Context window
Gemini 3.8 Flash
1M
GPT-6 Luna
1.1M
Maximum output
Gemini 3.8 Flash
66K
GPT-6 Luna
128K

Benchmarks & performance

Arena Elo
Gemini 3.8 Flash
1,244
GPT-6 Luna
1,211
MMLU-Pro
Gemini 3.8 Flash
81%
GPT-6 Luna
75.4%
Intelligence Index
Gemini 3.8 Flash
40.9
GPT-6 Luna
38.1
GPQA
Gemini 3.8 Flash
95.3%
GPT-6 Luna
Not listed
Humanity’s Last Exam
Gemini 3.8 Flash
47.8%
GPT-6 Luna
38.5%
SciCode
Gemini 3.8 Flash
56.6%
GPT-6 Luna
54.6%
Output speed
Gemini 3.8 Flash
239 tokens/s
GPT-6 Luna
141 tokens/s
Time to first token
Gemini 3.8 Flash
13.45 s
GPT-6 Luna
0.1 s
DesignArena rating
Gemini 3.8 Flash
1,308
GPT-6 Luna
1,276
DesignArena win rate
Gemini 3.8 Flash
50.3%
GPT-6 Luna
48.4%

Tools & features

Function calling
Gemini 3.8 Flash
Supported
GPT-6 Luna
Supported
Structured outputs
Gemini 3.8 Flash
Supported
GPT-6 Luna
Supported
JSON mode
Gemini 3.8 Flash
Supported
GPT-6 Luna
Supported
Reasoning
Gemini 3.8 Flash
Supported
GPT-6 Luna
Supported
Deterministic seed
Gemini 3.8 Flash
Supported
GPT-6 Luna
Supported
Prompt caching
Gemini 3.8 Flash
Supported
GPT-6 Luna
Supported

API & availability

API identifier
Gemini 3.8 Flash
gemini-3.8-flash
GPT-6 Luna
gpt-6-luna
API providers
Gemini 3.8 Flash
Google AI Studio, Google
GPT-6 Luna
OpenAI, Azure, Amazon Bedrock
Prices checked
Gemini 3.8 Flash
Oct 9, 2026
GPT-6 Luna
Oct 9, 2026

About the models

Gemini 3.8 Flash

Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows.

Full pricing & details

GPT-6 Luna

Smallest, fastest GPT-6 model for high-volume classification, extraction and chat. Long-context requests bill at $0.20 input / $0.75 output per 1M.

Full pricing & details

Comparison FAQ

Gemini 3.8 Flash: $0.75 input and $3.75 output per million tokens. GPT-6 Luna: $0.10 input and $0.50 output per million tokens. The cheaper choice depends on your input-to-output ratio; a model can have a lower input rate but a higher output rate.

For an example workload of 20 million input tokens and 5 million output tokens per month: Gemini 3.8 Flash: $34. GPT-6 Luna: $4.50. These are token-only estimates using the displayed rates, without caching discounts, additional tool charges or taxes. Actual bills depend on usage and endpoint pricing.

Gemini 3.8 Flash: 1M tokens of context; maximum output 66K tokens. GPT-6 Luna: 1.1M tokens of context; maximum output 128K tokens. Context capacity and maximum output are separate limits. A larger context window allows more material in a request but does not guarantee more accurate answers.

Gemini 3.8 Flash: Artificial Analysis SciCode 56.6%. GPT-6 Luna: Artificial Analysis SciCode 54.6%. Compare coding results within the same evaluation, then test representative tasks from your own codebase. A missing score is not a zero, and a single benchmark does not establish the best model for every programming task.

These evaluations measure different tasks and use different scales. Arena Elo and DesignArena ratings must remain separate, as must Artificial Analysis Intelligence Index and MMLU-Pro. Compare models within the same evaluation rather than combining scores. Missing results are marked Not listed, not treated as measured zeroes.

Gemini 3.8 Flash: Artificial Analysis output speed 239 tokens/s; time to first token 13.45 seconds. GPT-6 Luna: Artificial Analysis output speed 141 tokens/s; time to first token 0.1 seconds. Output speed measures generation throughput, while time to first token measures the initial wait. Real-world latency also depends on the API provider, request size and load; missing measurements cannot establish a speed winner.

Gemini 3.8 Flash: listed input formats: text, image, video, file, audio; output formats: text. GPT-6 Luna: listed input formats: file, image, text; output formats: text. Input and output support are different capabilities: accepting an image does not mean a model generates images. Confirm formats and file limits for the endpoint you plan to use.

Gemini 3.8 Flash: listed features: Function calling, Structured outputs, JSON mode, Reasoning, Deterministic seed, Prompt caching. GPT-6 Luna: listed features: Function calling, Structured outputs, JSON mode, Reasoning, Deterministic seed, Prompt caching. Only recorded capabilities are shown; an unlisted feature is not proof that it is unsupported. Availability can vary by API provider, so confirm tool calling and JSON schema support for your chosen endpoint.

Gemini 3.8 Flash: cached input price $0.07 per million tokens. GPT-6 Luna: cached input price $0.01 per million tokens. Caching may reduce charges for eligible repeated input when supported by the endpoint. A listed cached rate does not guarantee every request qualifies; cache duration, minimum token counts and any write charges depend on the provider.

Gemini 3.8 Flash: catalog status active; API identifier gemini-3.8-flash; listed API hosts Google AI Studio, Google. GPT-6 Luna: catalog status active; API identifier gpt-6-luna; listed API hosts OpenAI, Azure, Amazon Bedrock. Catalog status is not a live availability check. Before integrating, verify the provider's documentation, endpoint access, rate limits, regional availability and data-handling terms.

Popular model comparisons