Model pricing reference · Verified September 18, 2026

Gemini 3.1 Pro API Pricing: $2.00 / $12.00 per 1M Tokens

By AI Agent Hub Editorial Desk · Review method · Corrections

The numbers: Gemini 3.1 Pro (Google DeepMind) bills $2.00 per 1M input tokens, $12.00 per 1M output tokens, and $0.20 per 1M cached input tokens, inside a 1M context window. On our reference 200-step agent workload that works out to about $2.79 per task.

Google flagship for reasoning, agentic work, and long context; multimodal input. Every figure on this page was read directly from Google DeepMind’s own documentation on September 18, 2026 — not copied from a third-party comparison table. Rates change; the source link and check date below are part of the data.

Verified rates

ItemRateUnit
Input$2.00per 1M tokens
Output$12.00per 1M tokens
Cached input (cache read)$0.20per 1M tokens
Context window1Mtokens

Billing notes: Prompts above 200K input tokens bill at $4.00 input and $18.00 output per 1M for the whole request. Paid tier only since April 1, 2026; the free tier no longer covers Pro models.

Source: Google DeepMind official pricing ↗, checked September 18, 2026.

What a real workload costs on Gemini 3.1 Pro

List rates are not bills. These two reference workloads translate the table above into money. The assumptions are visible so you can swap in your own trace — the cost calculator does exactly that.

Scenario A — one 200-step agent task

A coding agent runs 200 steps against a stable 40,000-token prefix (system prompt, tools, repository map). The prefix is written to cache on the first step, re-read on the remaining 199, and each step produces 500 output tokens. For Google DeepMind, cache writes bill at the input rate (per-hour cache storage is excluded).

ComponentCalculationCost
Cache write (once)billed via hourly cache storage (excluded)—
Cache reads199 × 40,000 × $0.20 / 1M$1.59
Output200 × 500 × $12.00 / 1M$1.20
Total per task$2.79

Scenario B — one heavy month

A production service pushes 50M fresh input tokens, 150M cached input tokens, and 10M output tokens through Gemini 3.1 Pro in a month:

ComponentVolumeCost
Fresh input50M × $2.00$100.00
Cached input150M × $0.20$30.00
Output10M × $12.00$120.00
Monthly total$250.00

Head-to-head cost comparisons

Frequently asked questions

How much does Gemini 3.1 Pro cost per million tokens?

Gemini 3.1 Pro costs $2.00 per 1M input tokens and $12.00 per 1M output tokens, with cached input at $0.20 per 1M. Verified September 18, 2026 against Google DeepMind's official pricing.

What is the context window of Gemini 3.1 Pro?

Gemini 3.1 Pro supports a 1M token context window, per Google DeepMind's documentation checked September 18, 2026.

Does Gemini 3.1 Pro support prompt caching?

Yes. Cached input tokens bill at $0.20 per 1M instead of the full input rate; cache writes bill at the input rate (per-hour cache storage is excluded).

When was this Gemini 3.1 Pro price last checked?

On September 18, 2026, against Google DeepMind's own pricing page: https://deepmind.google/models/gemini/pro/

Keep digging