Coding Benchmark Matrix
A small comparable set of vendor-published SWE-Bench Pro and Terminal-Bench 2.1 results. This is not an independent AI Agent Hub laboratory test.
| Developer | Context | Input / output |
|---|
Benchmark results are only one input to a model decision. Tool configuration, inference settings, latency, reliability, and your own representative tasks matter. See the benchmark methodology.
How to read this matrix
The rows are limited to results published together with enough shared context to support a narrow comparison. A higher percentage means more tasks met that benchmark's scoring criteria in the cited run; it does not prove that the model will perform better in your repository, language, tool environment, or review process.
Comparable does not mean independent
The values are vendor-published. AI Agent Hub did not rerun the benchmark, verify every patch, or audit the evaluation harness.
Configuration matters
Agent scaffolding, reasoning budget, tool access, retries, model version, and inference settings can materially change a score.
Cost is incomplete here
The displayed unit prices do not include how many calls or tokens the benchmark run consumed. Do not infer cost per solved issue from this table.
Use benchmarks to form a shortlist
- Filter to models that meet your context and price constraints.
- Read the cited source and confirm the exact benchmark version and setup.
- Test the shortlist on representative tasks from your own workflow.
- Measure accepted outcomes, retries, latency, tool steps, and human review time.
Common interpretation mistakes
- Comparing scores from different benchmark versions as if they used the same task set.
- Treating a vendor result as an independent laboratory finding.
- Ignoring contamination, harness changes, or missing confidence intervals.
- Choosing a model from one score without testing safety and operational reliability.