AI API Cost Calculator

Estimate workload cost using the current source-linked model catalog. Change every material assumption below; this is not a provider bill or subscription quote.

Workload assumptions

Estimated monthly API usage
$0.00


Input cost:

Output cost:

Estimated cache savings:

Same workload across verified models

ModelPublished input / outputEstimated monthDifference
Excluded: taxes, currency conversion, account tiers, reseller markup, batch discounts, minimum commitments, and charges outside token usage. Cache math uses each displayed cache rate. Where a provider publishes no cache-read price, no discount is assumed. Read the calculator methodology.

How to build a useful estimate

Start with usage from a representative trace rather than a guessed “average prompt.” Include the system message, retrieved documents, tool results, conversation history, and expected output. Calls per day should count every model request inside an agent workflow—not only the number of user sessions.

1. Model the typical case

Use median input and output sizes for routine traffic. Set cache share to zero unless your provider usage metadata confirms repeated-prefix hits.

2. Stress the expensive case

Repeat the estimate with long inputs, longer outputs, retries, and more working days. This gives a range instead of a misleading single forecast.

3. Compare accepted outcomes

A lower per-call estimate is not automatically cheaper if the model needs more retries or tool steps. Compare cost per task that passes your evaluation.

What the calculator actually computes

For each model, the tool multiplies monthly uncached input tokens by the published input rate, cached input tokens by the displayed cache-read rate, and output tokens by the output rate. It then compares the same workload across the reviewed catalog. It does not predict future prices or provider routing behavior.

Questions this result cannot answer

Frequently asked questions

No. The calculator applies provider-published rates per one million tokens to workload assumptions you choose. It excludes taxes, currency conversion, account tiers, reseller markup, batch discounts, minimum commitments, and any charge outside token usage.
Three causes dominate. Retries and failed calls are still billed. Agent workflows often make several model calls per single user action, so counting user sessions undercounts requests. Retrieved documents, tool results, and conversation history are all billable input tokens, not just the newest message.
Every rate comes from the catalog snapshot stated at the top of this page, and each model row links to its official provider source. Prices in this category change several times a year, so check the source link before making a material commitment.
Both. Run the routine median first, then repeat it with long inputs, longer outputs, more retries, and more working days. A range is more useful than a single number, and the upper bound is what protects a budget.