Coding Benchmark Matrix

A small comparable set of vendor-published SWE-Bench Pro and Terminal-Bench 2.1 results. This is not an independent AI Agent Hub laboratory test.

DeveloperContextInput / output

Benchmark results are only one input to a model decision. Tool configuration, inference settings, latency, reliability, and your own representative tasks matter. See the benchmark methodology.

How to read this matrix

The rows are limited to results published together with enough shared context to support a narrow comparison. A higher percentage means more tasks met that benchmark's scoring criteria in the cited run; it does not prove that the model will perform better in your repository, language, tool environment, or review process.

Comparable does not mean independent

The values are vendor-published. AI Agent Hub did not rerun the benchmark, verify every patch, or audit the evaluation harness.

Configuration matters

Agent scaffolding, reasoning budget, tool access, retries, model version, and inference settings can materially change a score.

Cost is incomplete here

The displayed unit prices do not include how many calls or tokens the benchmark run consumed. Do not infer cost per solved issue from this table.

Use benchmarks to form a shortlist

  1. Filter to models that meet your context and price constraints.
  2. Read the cited source and confirm the exact benchmark version and setup.
  3. Test the shortlist on representative tasks from your own workflow.
  4. Measure accepted outcomes, retries, latency, tool steps, and human review time.

Common interpretation mistakes

Frequently asked questions

No. The matrix is a small comparable set of vendor-published results, and the page says so explicitly. Rows are limited to results published with enough shared context to support a narrow comparison; they are not presented as independent laboratory measurements.
Published runs use a specific harness, tool configuration, inference settings, and task set. Your repository, language mix, review process, and latency constraints differ. A benchmark score is evidence about one setting, not a prediction about yours.
The one closest to your work, and no single one should decide it. SWE-bench-style results say most about repository-scale coding tasks; terminal and multimodal results measure different capabilities. Treat all of them as one input alongside your own evaluation.
It means more tasks met that benchmark's scoring criteria in the cited run. It does not prove the model will perform better in your environment, and it says nothing about cost, latency, or how the model fails.