Head-to-head cost comparison · Verified September 18, 2026

GPT-5.6 Luna vs Gemini 3.8 Flash: API Cost Compared

By AI Agent Hub Editorial Desk · Review method · Corrections

The verdict, in numbers: On our reference 200-step agent workload, GPT-5.6 Luna bills $0.29 against $0.97 for Gemini 3.8 Flash — 3.4x apart. List rates alone would not have told you that.

This page compares list rates and workload cost only. It does not tell you which model is better — capability belongs to your own evaluation. What it does give you is the half of the decision that can be verified: what each provider will actually charge, read from their own documentation on September 18, 2026.

Rate cards side by side

ItemGPT-5.6 LunaGemini 3.8 Flash
ProviderOpenAIGoogle DeepMind
Input / 1M$0.20$0.75
Output / 1M$1.20$3.75
Cache read / 1M$0.020$0.075
Context window1.05M1M

Sources: OpenAI pricing ↗ and Google DeepMind pricing ↗, both checked September 18, 2026.

The same 200-step agent task, priced twice

One coding-agent task: 200 steps, a stable 40,000-token prefix written to cache on step 1 and re-read on the remaining 199, 500 output tokens per step. Cache-write billing follows each provider’s own rules (cache writes bill at 1.25x the uncached input rate; cache writes bill at the input rate (per-hour cache storage is excluded)).

ComponentGPT-5.6 LunaGemini 3.8 Flash
Cache write (once)$0.01—
Cache reads (199 × 40K)$0.16$0.60
Output (200 × 500)$0.12$0.38
Total per task$0.29$0.97

A heavy month on each

50M fresh input, 150M cached input, 10M output tokens in a month:

ComponentGPT-5.6 LunaGemini 3.8 Flash
Fresh input (50M)$10.00$37.50
Cached input (150M)$3.00$11.25
Output (10M)$12.00$37.50
Monthly total$25.00$86.25

What this comparison does and does not say

GPT-5.6 Luna: Lowest-cost GPT-5.6 tier; cheapest current-generation frontier rate published by OpenAI.

Gemini 3.8 Flash: Newest Flash model for long-horizon software engineering and autonomous agents; 1M context, 64K output.

Cost is the verifiable half of the decision; quality is the half only your workload can answer. Run both candidates on a representative evaluation set and measure cost per accepted task — the cost planning guide walks through that method, and the calculator lets you re-price this exact comparison with your own token mix.

Frequently asked questions

Is GPT-5.6 Luna cheaper than Gemini 3.8 Flash?

On list rates, GPT-5.6 Luna costs $0.20/$1.20 per 1M input/output tokens versus $0.75/$3.75 for Gemini 3.8 Flash. On a 200-step agent workload (40K cached prefix, 500 output tokens per step), GPT-5.6 Luna totals $0.29 versus $0.97 for Gemini 3.8 Flash. Prices verified September 18, 2026.

GPT-5.6 Luna vs Gemini 3.8 Flash: which has the longer context window?

GPT-5.6 Luna offers 1.05M tokens; Gemini 3.8 Flash offers 1M tokens, per provider documentation checked September 18, 2026.

Where do these GPT-5.6 Luna and Gemini 3.8 Flash prices come from?

Both rate cards were read from the providers' own pricing pages on September 18, 2026: https://developers.openai.com/api/docs/models/gpt-5.6-luna and https://deepmind.google/models/gemini/flash/.

Keep digging