Models, measured side by side
Gemini 3.8 Flash vs GPT-6 Luna
Compare Gemini 3.8 Flash, GPT-6 Luna by token pricing, context length, benchmark results, speed and tool support.
Updated Oct 2026 · Last verified Oct 9, 2026 · Public listed prices. Confirm with the provider.
Verdict
Gemini 3.8 Flash vs GPT-6 LunaCheaper
GPT-6 Luna
87% less than next at 3:1
Smarter
Gemini 3.8 Flash
AA 40.9 vs 38.1
Bigger context
GPT-6 Luna
1.1M vs 1M
Choose Gemini 3.8 Flash if…
- you need the strongest reasoning (AA Intelligence 40.9 vs 38.1)
Choose GPT-6 Luna if…
- you want the lowest cost on input-heavy work (87% cheaper at 3:1)
- you generate a lot of output ($0.500 vs $3.75 per 1M output tokens)
- you need the biggest context window (1.1M vs 1M tokens)
10M input + 2M output / monthGemini 3.8 Flash: $15.00GPT-6 Luna: $2.00
Monthly workload
Monthly cost breakdown
InputOutput
Gemini 3.8 Flash
$34 / mo
GPT-6 Luna
$4.50 / mo
Performance comparison
Gemini 3.8 Flash
40.9
GPT-6 Luna
38.1
Overview
Pricing & capacity
| Pricing & capacity | Gemini 3.8 Flash | GPT-6 Luna |
|---|---|---|
| Input / 1M tokens | $0.75 | $0.10 |
| Output / 1M tokens | $3.75 | $0.50 |
| Cached input / 1M | $0.07 | $0.01 |
| Monthly cost | $34 | $4.50 |
| Context window | 1M | 1.1M |
| Maximum output | 66K | 128K |
Input / 1M tokens
- Gemini 3.8 Flash
- $0.75
- GPT-6 Luna
- $0.10
Output / 1M tokens
- Gemini 3.8 Flash
- $3.75
- GPT-6 Luna
- $0.50
Cached input / 1M
- Gemini 3.8 Flash
- $0.07
- GPT-6 Luna
- $0.01
Monthly cost
- Gemini 3.8 Flash
- $34
- GPT-6 Luna
- $4.50
Context window
- Gemini 3.8 Flash
- 1M
- GPT-6 Luna
- 1.1M
Maximum output
- Gemini 3.8 Flash
- 66K
- GPT-6 Luna
- 128K
Benchmarks & performance
| Benchmarks & performance | Gemini 3.8 Flash | GPT-6 Luna |
|---|---|---|
| Arena Elo | 1,244 | 1,211 |
| MMLU-Pro | 81% | 75.4% |
| Intelligence Index | 40.9 | 38.1 |
| GPQA | 95.3% | Not listed |
| Humanity’s Last Exam | 47.8% | 38.5% |
| SciCode | 56.6% | 54.6% |
| Output speed | 239 tokens/s | 141 tokens/s |
| Time to first token | 13.45 s | 0.1 s |
| DesignArena rating | 1,308 | 1,276 |
| DesignArena win rate | 50.3% | 48.4% |
Arena Elo
- Gemini 3.8 Flash
- 1,244
- GPT-6 Luna
- 1,211
MMLU-Pro
- Gemini 3.8 Flash
- 81%
- GPT-6 Luna
- 75.4%
Intelligence Index
- Gemini 3.8 Flash
- 40.9
- GPT-6 Luna
- 38.1
GPQA
- Gemini 3.8 Flash
- 95.3%
- GPT-6 Luna
- Not listed
Humanity’s Last Exam
- Gemini 3.8 Flash
- 47.8%
- GPT-6 Luna
- 38.5%
SciCode
- Gemini 3.8 Flash
- 56.6%
- GPT-6 Luna
- 54.6%
Output speed
- Gemini 3.8 Flash
- 239 tokens/s
- GPT-6 Luna
- 141 tokens/s
Time to first token
- Gemini 3.8 Flash
- 13.45 s
- GPT-6 Luna
- 0.1 s
DesignArena rating
- Gemini 3.8 Flash
- 1,308
- GPT-6 Luna
- 1,276
DesignArena win rate
- Gemini 3.8 Flash
- 50.3%
- GPT-6 Luna
- 48.4%
Tools & features
| Tools & features | Gemini 3.8 Flash | GPT-6 Luna |
|---|---|---|
| Function calling | Supported | Supported |
| Structured outputs | Supported | Supported |
| JSON mode | Supported | Supported |
| Reasoning | Supported | Supported |
| Deterministic seed | Supported | Supported |
| Prompt caching | Supported | Supported |
Function calling
- Gemini 3.8 Flash
- Supported
- GPT-6 Luna
- Supported
Structured outputs
- Gemini 3.8 Flash
- Supported
- GPT-6 Luna
- Supported
JSON mode
- Gemini 3.8 Flash
- Supported
- GPT-6 Luna
- Supported
Reasoning
- Gemini 3.8 Flash
- Supported
- GPT-6 Luna
- Supported
Deterministic seed
- Gemini 3.8 Flash
- Supported
- GPT-6 Luna
- Supported
Prompt caching
- Gemini 3.8 Flash
- Supported
- GPT-6 Luna
- Supported
API & availability
| API & availability | Gemini 3.8 Flash | GPT-6 Luna |
|---|---|---|
| API identifier | gemini-3.8-flash | gpt-6-luna |
| API providers | Google AI Studio, Google | OpenAI, Azure, Amazon Bedrock |
| Prices checked | Oct 9, 2026 | Oct 9, 2026 |
API identifier
- Gemini 3.8 Flash
gemini-3.8-flash- GPT-6 Luna
gpt-6-luna
API providers
- Gemini 3.8 Flash
- Google AI Studio, Google
- GPT-6 Luna
- OpenAI, Azure, Amazon Bedrock
Prices checked
- Gemini 3.8 Flash
- Oct 9, 2026
- GPT-6 Luna
- Oct 9, 2026
About the models
Gemini 3.8 Flash
Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows.
Full pricing & detailsGPT-6 Luna
Smallest, fastest GPT-6 model for high-volume classification, extraction and chat. Long-context requests bill at $0.20 input / $0.75 output per 1M.
Full pricing & detailsComparison FAQ
Gemini 3.8 Flash: $0.75 input and $3.75 output per million tokens. GPT-6 Luna: $0.10 input and $0.50 output per million tokens. The cheaper choice depends on your input-to-output ratio; a model can have a lower input rate but a higher output rate.
For an example workload of 20 million input tokens and 5 million output tokens per month: Gemini 3.8 Flash: $34. GPT-6 Luna: $4.50. These are token-only estimates using the displayed rates, without caching discounts, additional tool charges or taxes. Actual bills depend on usage and endpoint pricing.
Gemini 3.8 Flash: 1M tokens of context; maximum output 66K tokens. GPT-6 Luna: 1.1M tokens of context; maximum output 128K tokens. Context capacity and maximum output are separate limits. A larger context window allows more material in a request but does not guarantee more accurate answers.
Gemini 3.8 Flash: Artificial Analysis SciCode 56.6%. GPT-6 Luna: Artificial Analysis SciCode 54.6%. Compare coding results within the same evaluation, then test representative tasks from your own codebase. A missing score is not a zero, and a single benchmark does not establish the best model for every programming task.
These evaluations measure different tasks and use different scales. Arena Elo and DesignArena ratings must remain separate, as must Artificial Analysis Intelligence Index and MMLU-Pro. Compare models within the same evaluation rather than combining scores. Missing results are marked Not listed, not treated as measured zeroes.
Gemini 3.8 Flash: Artificial Analysis output speed 239 tokens/s; time to first token 13.45 seconds. GPT-6 Luna: Artificial Analysis output speed 141 tokens/s; time to first token 0.1 seconds. Output speed measures generation throughput, while time to first token measures the initial wait. Real-world latency also depends on the API provider, request size and load; missing measurements cannot establish a speed winner.
Gemini 3.8 Flash: listed input formats: text, image, video, file, audio; output formats: text. GPT-6 Luna: listed input formats: file, image, text; output formats: text. Input and output support are different capabilities: accepting an image does not mean a model generates images. Confirm formats and file limits for the endpoint you plan to use.
Gemini 3.8 Flash: listed features: Function calling, Structured outputs, JSON mode, Reasoning, Deterministic seed, Prompt caching. GPT-6 Luna: listed features: Function calling, Structured outputs, JSON mode, Reasoning, Deterministic seed, Prompt caching. Only recorded capabilities are shown; an unlisted feature is not proof that it is unsupported. Availability can vary by API provider, so confirm tool calling and JSON schema support for your chosen endpoint.
Gemini 3.8 Flash: cached input price $0.07 per million tokens. GPT-6 Luna: cached input price $0.01 per million tokens. Caching may reduce charges for eligible repeated input when supported by the endpoint. A listed cached rate does not guarantee every request qualifies; cache duration, minimum token counts and any write charges depend on the provider.
Gemini 3.8 Flash: catalog status active; API identifier gemini-3.8-flash; listed API hosts Google AI Studio, Google. GPT-6 Luna: catalog status active; API identifier gpt-6-luna; listed API hosts OpenAI, Azure, Amazon Bedrock. Catalog status is not a live availability check. Before integrating, verify the provider's documentation, endpoint access, rate limits, regional availability and data-handling terms.
