GPT-5.6 Luna vs Gemini 3.8 Flash: API Cost Compared
The verdict, in numbers: On our reference 200-step agent workload, GPT-5.6 Luna bills $0.29 against $0.97 for Gemini 3.8 Flash — 3.4x apart. List rates alone would not have told you that.
This page compares list rates and workload cost only. It does not tell you which model is better — capability belongs to your own evaluation. What it does give you is the half of the decision that can be verified: what each provider will actually charge, read from their own documentation on September 18, 2026.
Rate cards side by side
| Item | GPT-5.6 Luna | Gemini 3.8 Flash |
|---|---|---|
| Provider | OpenAI | Google DeepMind |
| Input / 1M | $0.20 | $0.75 |
| Output / 1M | $1.20 | $3.75 |
| Cache read / 1M | $0.020 | $0.075 |
| Context window | 1.05M | 1M |
Sources: OpenAI pricing ↗ and Google DeepMind pricing ↗, both checked September 18, 2026.
The same 200-step agent task, priced twice
One coding-agent task: 200 steps, a stable 40,000-token prefix written to cache on step 1 and re-read on the remaining 199, 500 output tokens per step. Cache-write billing follows each provider’s own rules (cache writes bill at 1.25x the uncached input rate; cache writes bill at the input rate (per-hour cache storage is excluded)).
| Component | GPT-5.6 Luna | Gemini 3.8 Flash |
|---|---|---|
| Cache write (once) | $0.01 | — |
| Cache reads (199 × 40K) | $0.16 | $0.60 |
| Output (200 × 500) | $0.12 | $0.38 |
| Total per task | $0.29 | $0.97 |
A heavy month on each
50M fresh input, 150M cached input, 10M output tokens in a month:
| Component | GPT-5.6 Luna | Gemini 3.8 Flash |
|---|---|---|
| Fresh input (50M) | $10.00 | $37.50 |
| Cached input (150M) | $3.00 | $11.25 |
| Output (10M) | $12.00 | $37.50 |
| Monthly total | $25.00 | $86.25 |
What this comparison does and does not say
GPT-5.6 Luna: Lowest-cost GPT-5.6 tier; cheapest current-generation frontier rate published by OpenAI.
Gemini 3.8 Flash: Newest Flash model for long-horizon software engineering and autonomous agents; 1M context, 64K output.
Cost is the verifiable half of the decision; quality is the half only your workload can answer. Run both candidates on a representative evaluation set and measure cost per accepted task — the cost planning guide walks through that method, and the calculator lets you re-price this exact comparison with your own token mix.
Frequently asked questions
Is GPT-5.6 Luna cheaper than Gemini 3.8 Flash?
On list rates, GPT-5.6 Luna costs $0.20/$1.20 per 1M input/output tokens versus $0.75/$3.75 for Gemini 3.8 Flash. On a 200-step agent workload (40K cached prefix, 500 output tokens per step), GPT-5.6 Luna totals $0.29 versus $0.97 for Gemini 3.8 Flash. Prices verified September 18, 2026.
GPT-5.6 Luna vs Gemini 3.8 Flash: which has the longer context window?
GPT-5.6 Luna offers 1.05M tokens; Gemini 3.8 Flash offers 1M tokens, per provider documentation checked September 18, 2026.
Where do these GPT-5.6 Luna and Gemini 3.8 Flash prices come from?
Both rate cards were read from the providers' own pricing pages on September 18, 2026: https://developers.openai.com/api/docs/models/gpt-5.6-luna and https://deepmind.google/models/gemini/flash/.