Interactive AI Model & Agent Tools
Five focused utilities use source-linked data or generate transparent design artifacts. Calculators expose assumptions; design tools do not claim to replace evaluation or production review.
Workload calculator
Adjust model, calls, token volume, cache share, and working days.
Open calculator →
Text estimator
Estimate text size and preview one call across the verified catalog.
Open estimator →
Attributed matrix
Filter a comparable vendor-published set with source and caveat.
Open matrix →
Rule-based decision aid
Create a three-model shortlist by workload, unit-cost priority, and context need.
Build shortlist →
Architecture worksheet
Generate a reviewable JSON blueprint covering permissions, approvals, limits, and evaluation.
Build blueprint →
How to choose between these tools
Each utility answers a different question, and using the wrong one produces a number that looks precise but answers nothing. Start from the decision you actually need to make rather than from the tool that sounds most impressive.
- You have a workload and need a budget. Use the API cost calculator. It multiplies your own call volume, token mix, and cache share against published per-million-token rates, so the output is only as good as the traffic sample you enter.
- You have a single prompt and want its size. Use the token counter. It estimates tokens for pasted text and previews one request across the verified catalog. Treat the estimate as an approximation: tokenisation differs by provider and, in Anthropic's case, changed with the Claude 4.7 tokenizer.
- You are comparing published results. Use the benchmark matrix. It deliberately carries a small comparable set rather than a long leaderboard, because vendor-reported scores often differ in harness, tool access, and inference settings.
- You do not yet know which model to evaluate. Use the shortlist wizard. It applies explicit rules about workload, unit-price priority, and context need, then shows the reasoning so you can disagree with it.
- You are designing an agent, not picking a model. Use the blueprint builder. It produces a reviewable artifact covering permissions, approvals, spending limits, and evaluation rather than a score.
Where the numbers come from
Every rate in these tools is read from a single shared catalog maintained at js/model-data.js, and each entry carries the provider documentation URL and the date it was checked. Prices are copied from the provider's own pricing or model page, never from a third-party comparison table. Where a provider publishes a promotional rate, the expiry date travels with the number. Where a provider separates peak from off-peak billing, the tool shows the off-peak rate and states the peak multiplier beside it.
Limits worth understanding before you trust an output
- A unit-price comparison is not a cost comparison. A cheaper input rate loses to a model that finishes in one call instead of five, retries less, or needs fewer agent steps.
- Long-context billing is not always proportional. Several providers charge a premium across the whole request once input crosses a threshold, rather than only on the excess. The catalog notes where this applies.
- Cache behaviour is the largest variable. A cache hit can cut input cost by an order of magnitude, but only when a stable prefix is genuinely reused. Measure your own hit rate instead of assuming one.
- Cache writes bill too. Writing to cache is not free on every provider, so a workload with a low hit rate can cost more with caching enabled than without.
- Benchmark scores are evidence about one setting. They are not predictions about your repository, language mix, or review process.
- Rates move. The catalog is re-verified on a schedule and each page shows its verification date. Check that date before quoting a figure in a budget.
These tools run entirely in your browser. Text you paste is not uploaded, and no entered value is stored or transmitted. If an output leads to a decision with real cost attached, confirm the rate against the provider's pricing page on the day you commit.