Head-to-head cost comparison · Verified September 18, 2026

Claude Haiku 4.5 vs Gemini 3.8 Flash: API Cost Compared

By AI Agent Hub Editorial Desk · Review method · Corrections

The verdict, in numbers: On our reference 200-step agent workload, Gemini 3.8 Flash bills $0.97 against $1.35 for Claude Haiku 4.5 — 1.4x apart. List rates alone would not have told you that.

This page compares list rates and workload cost only. It does not tell you which model is better — capability belongs to your own evaluation. What it does give you is the half of the decision that can be verified: what each provider will actually charge, read from their own documentation on September 18, 2026.

Rate cards side by side

ItemClaude Haiku 4.5Gemini 3.8 Flash
ProviderAnthropicGoogle DeepMind
Input / 1M$1.00$0.75
Output / 1M$5.00$3.75
Cache read / 1M$0.10$0.075
Context window200K1M

Sources: Anthropic pricing ↗ and Google DeepMind pricing ↗, both checked September 18, 2026.

The same 200-step agent task, priced twice

One coding-agent task: 200 steps, a stable 40,000-token prefix written to cache on step 1 and re-read on the remaining 199, 500 output tokens per step. Cache-write billing follows each provider’s own rules (cache writes bill at 1.25x the input rate; cache writes bill at the input rate (per-hour cache storage is excluded)).

ComponentClaude Haiku 4.5Gemini 3.8 Flash
Cache write (once)$0.05—
Cache reads (199 × 40K)$0.80$0.60
Output (200 × 500)$0.50$0.38
Total per task$1.35$0.97

A heavy month on each

50M fresh input, 150M cached input, 10M output tokens in a month:

ComponentClaude Haiku 4.5Gemini 3.8 Flash
Fresh input (50M)$50.00$37.50
Cached input (150M)$15.00$11.25
Output (10M)$50.00$37.50
Monthly total$115.00$86.25

What this comparison does and does not say

Claude Haiku 4.5: Fast current Claude model for high-volume and latency-sensitive work.

Gemini 3.8 Flash: Newest Flash model for long-horizon software engineering and autonomous agents; 1M context, 64K output.

Cost is the verifiable half of the decision; quality is the half only your workload can answer. Run both candidates on a representative evaluation set and measure cost per accepted task — the cost planning guide walks through that method, and the calculator lets you re-price this exact comparison with your own token mix.

Frequently asked questions

Is Claude Haiku 4.5 cheaper than Gemini 3.8 Flash?

On list rates, Claude Haiku 4.5 costs $1.00/$5.00 per 1M input/output tokens versus $0.75/$3.75 for Gemini 3.8 Flash. On a 200-step agent workload (40K cached prefix, 500 output tokens per step), Gemini 3.8 Flash totals $0.97 versus $1.35 for Claude Haiku 4.5. Prices verified September 18, 2026.

Claude Haiku 4.5 vs Gemini 3.8 Flash: which has the longer context window?

Claude Haiku 4.5 offers 200K tokens; Gemini 3.8 Flash offers 1M tokens, per provider documentation checked September 18, 2026.

Where do these Claude Haiku 4.5 and Gemini 3.8 Flash prices come from?

Both rate cards were read from the providers' own pricing pages on September 18, 2026: https://platform.claude.com/docs/en/about-claude/pricing and https://deepmind.google/models/gemini/flash/.

Keep digging