AI API Cost Calculator
Estimate workload cost using the current source-linked model catalog. Change every material assumption below; this is not a provider bill or subscription quote.
Workload assumptions
Input cost:
Output cost:
Estimated cache savings:
Same workload across verified models
| Model | Published input / output | Estimated month | Difference |
|---|
How to build a useful estimate
Start with usage from a representative trace rather than a guessed “average prompt.” Include the system message, retrieved documents, tool results, conversation history, and expected output. Calls per day should count every model request inside an agent workflow—not only the number of user sessions.
1. Model the typical case
Use median input and output sizes for routine traffic. Set cache share to zero unless your provider usage metadata confirms repeated-prefix hits.
2. Stress the expensive case
Repeat the estimate with long inputs, longer outputs, retries, and more working days. This gives a range instead of a misleading single forecast.
3. Compare accepted outcomes
A lower per-call estimate is not automatically cheaper if the model needs more retries or tool steps. Compare cost per task that passes your evaluation.
What the calculator actually computes
For each model, the tool multiplies monthly uncached input tokens by the published input rate, cached input tokens by the displayed cache-read rate, and output tokens by the output rate. It then compares the same workload across the reviewed catalog. It does not predict future prices or provider routing behavior.
Questions this result cannot answer
- Whether the selected model is accurate enough for your task.
- Whether your prompt is eligible for a provider cache or retention tier.
- How much retries, tools, search, embeddings, storage, or human review will cost.
- Whether a promotional rate will still apply when the workload launches.