AI Model Shortlist Wizard

Create a starting shortlist from the verified catalog. This is a transparent rule-based aid, not a sponsored ranking or a prediction of your monthly bill.

Data snapshot: . Final selection should include your own quality, latency, privacy, and reliability tests.

Primary workload
Primary priority
Context requirement

Your shortlist

Rates are USD per one million input/output tokens. Temporary pricing and cache assumptions are shown where applicable. Use the cost calculator after estimating your actual workload.

What this wizard does

The wizard applies visible rules to a deliberately small reviewed catalog. Priority selects one of three maintained candidate groups; workload can add models associated with complex, multimodal, or high-volume use; and the context choice removes models below the requested window. The result is a shortlist, not a winner.

Cost priority

Favors lower published input/output rates. It does not know your prompt length, cache-hit rate, retries, or quality threshold.

Balanced priority

Includes models positioned between unit cost and broad capability. “Balanced” is an editorial rule, not a scientific score.

Capability priority

Starts with models intended for more demanding work. You still need task-specific evidence before deployment.

How to validate the shortlist

  1. Write an acceptance test for the actual task.
  2. Run the same representative examples through every candidate.
  3. Record task success, critical failures, latency, token usage, retries, and tool calls.
  4. Reject candidates that fail privacy or safety gates, regardless of their weighted score.
  5. Choose by cost per accepted outcome and keep a rollback option.

Important limitations

Frequently asked questions

No. It is a transparent rule-based aid: your priority setting selects one of three maintained candidate groups, and your workload and context requirements narrow the list. There is no paid placement in the catalog.
No. It compares published unit rates only. Monthly spend depends on your call volume, input and output sizes, cache behaviour, and retry rate, which is what the cost calculator is for.
Only if it clears your quality bar. A cheaper model that needs more retries, more tool steps, or human correction can cost more per completed task. Compare cost per task that passes your evaluation, not cost per call.
Test the top two or three candidates on your own representative tasks, measuring quality, latency, and reliability under your real tool configuration. Then re-check the official source link for each rate before committing.