DeepSeek V4.1 Flash API Pricing: $0.15 / $0.60 per 1M Tokens
The numbers: DeepSeek V4.1 Flash (DeepSeek) bills $0.15 per 1M input tokens, $0.60 per 1M output tokens, and $0.0030 per 1M cached input tokens, inside a 1M context window. On our reference 200-step agent workload that works out to about $0.09 per task.
Low-cost current DeepSeek API model with thinking mode, vision, and Responses API support. Every figure on this page was read directly from DeepSeek’s own documentation on September 18, 2026 — not copied from a third-party comparison table. Rates change; the source link and check date below are part of the data.
Verified rates
| Item | Rate | Unit |
|---|---|---|
| Input | $0.15 | per 1M tokens |
| Output | $0.60 | per 1M tokens |
| Cached input (cache read) | $0.0030 | per 1M tokens |
| Context window | 1M | tokens |
Billing notes: Off-peak rate shown. Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, when input doubles to $0.30 and output to $1.20 per 1M. Model name is deepseek-flash; the retired deepseek-v4-flash identifier still resolves here.
Source: DeepSeek official pricing ↗, checked September 18, 2026.
What a real workload costs on DeepSeek V4.1 Flash
List rates are not bills. These two reference workloads translate the table above into money. The assumptions are visible so you can swap in your own trace — the cost calculator does exactly that.
Scenario A — one 200-step agent task
A coding agent runs 200 steps against a stable 40,000-token prefix (system prompt, tools, repository map). The prefix is written to cache on the first step, re-read on the remaining 199, and each step produces 500 output tokens. For DeepSeek, cache writes bill at the standard input rate.
| Component | Calculation | Cost |
|---|---|---|
| Cache write (once) | 40,000 tokens × $0.15 / 1M | $0.01 |
| Cache reads | 199 × 40,000 × $0.0030 / 1M | $0.02 |
| Output | 200 × 500 × $0.60 / 1M | $0.06 |
| Total per task | $0.09 |
Scenario B — one heavy month
A production service pushes 50M fresh input tokens, 150M cached input tokens, and 10M output tokens through DeepSeek V4.1 Flash in a month:
| Component | Volume | Cost |
|---|---|---|
| Fresh input | 50M × $0.15 | $7.50 |
| Cached input | 150M × $0.0030 | $0.45 |
| Output | 10M × $0.60 | $6.00 |
| Monthly total | $13.95 |
DeepSeek rates above are off-peak; weekday peak hours (01:00–04:00 and 06:00–10:00 UTC) double the input and output rates.
Head-to-head cost comparisons
- GPT-5.6 Luna vs DeepSeek V4.1 Flash — same workload, $0.29 vs $0.09 (3.2x apart).
- Gemini 3.8 Flash vs DeepSeek V4.1 Flash — same workload, $0.97 vs $0.09 (10.8x apart).
Frequently asked questions
How much does DeepSeek V4.1 Flash cost per million tokens?
DeepSeek V4.1 Flash costs $0.15 per 1M input tokens and $0.60 per 1M output tokens, with cached input at $0.0030 per 1M. Verified September 18, 2026 against DeepSeek's official pricing.
What is the context window of DeepSeek V4.1 Flash?
DeepSeek V4.1 Flash supports a 1M token context window, per DeepSeek's documentation checked September 18, 2026.
Does DeepSeek V4.1 Flash support prompt caching?
Yes. Cached input tokens bill at $0.0030 per 1M instead of the full input rate; cache writes bill at the standard input rate.
When was this DeepSeek V4.1 Flash price last checked?
On September 18, 2026, against DeepSeek's own pricing page: https://api-docs.deepseek.com/quick_start/pricing/