Gemini 3.8 Flash vs DeepSeek V4.1 Flash: API Cost Compared
The verdict, in numbers: On our reference 200-step agent workload, DeepSeek V4.1 Flash bills $0.09 against $0.97 for Gemini 3.8 Flash — 10.8x apart. List rates alone would not have told you that.
This page compares list rates and workload cost only. It does not tell you which model is better — capability belongs to your own evaluation. What it does give you is the half of the decision that can be verified: what each provider will actually charge, read from their own documentation on September 18, 2026.
Rate cards side by side
| Item | Gemini 3.8 Flash | DeepSeek V4.1 Flash |
|---|---|---|
| Provider | Google DeepMind | DeepSeek |
| Input / 1M | $0.75 | $0.15 |
| Output / 1M | $3.75 | $0.60 |
| Cache read / 1M | $0.075 | $0.0030 |
| Context window | 1M | 1M |
Sources: Google DeepMind pricing ↗ and DeepSeek pricing ↗, both checked September 18, 2026.
The same 200-step agent task, priced twice
One coding-agent task: 200 steps, a stable 40,000-token prefix written to cache on step 1 and re-read on the remaining 199, 500 output tokens per step. Cache-write billing follows each provider’s own rules (cache writes bill at the input rate (per-hour cache storage is excluded); cache writes bill at the standard input rate).
| Component | Gemini 3.8 Flash | DeepSeek V4.1 Flash |
|---|---|---|
| Cache write (once) | — | $0.01 |
| Cache reads (199 × 40K) | $0.60 | $0.02 |
| Output (200 × 500) | $0.38 | $0.06 |
| Total per task | $0.97 | $0.09 |
A heavy month on each
50M fresh input, 150M cached input, 10M output tokens in a month:
| Component | Gemini 3.8 Flash | DeepSeek V4.1 Flash |
|---|---|---|
| Fresh input (50M) | $37.50 | $7.50 |
| Cached input (150M) | $11.25 | $0.45 |
| Output (10M) | $37.50 | $6.00 |
| Monthly total | $86.25 | $13.95 |
DeepSeek figures use off-peak rates; weekday peak hours (01:00–04:00 and 06:00–10:00 UTC) double input and output.
What this comparison does and does not say
Gemini 3.8 Flash: Newest Flash model for long-horizon software engineering and autonomous agents; 1M context, 64K output.
DeepSeek V4.1 Flash: Low-cost current DeepSeek API model with thinking mode, vision, and Responses API support.
Cost is the verifiable half of the decision; quality is the half only your workload can answer. Run both candidates on a representative evaluation set and measure cost per accepted task — the cost planning guide walks through that method, and the calculator lets you re-price this exact comparison with your own token mix.
Frequently asked questions
Is Gemini 3.8 Flash cheaper than DeepSeek V4.1 Flash?
On list rates, Gemini 3.8 Flash costs $0.75/$3.75 per 1M input/output tokens versus $0.15/$0.60 for DeepSeek V4.1 Flash. On a 200-step agent workload (40K cached prefix, 500 output tokens per step), DeepSeek V4.1 Flash totals $0.09 versus $0.97 for Gemini 3.8 Flash. Prices verified September 18, 2026.
Gemini 3.8 Flash vs DeepSeek V4.1 Flash: which has the longer context window?
Gemini 3.8 Flash offers 1M tokens; DeepSeek V4.1 Flash offers 1M tokens, per provider documentation checked September 18, 2026.
Where do these Gemini 3.8 Flash and DeepSeek V4.1 Flash prices come from?
Both rate cards were read from the providers' own pricing pages on September 18, 2026: https://deepmind.google/models/gemini/flash/ and https://api-docs.deepseek.com/quick_start/pricing/.