Deep Dive into SWE-bench Verified & Terminal-Bench v2

By AI Agent Hub Editorial Desk · Review method · Corrections

Benchmarking · 5 min read · Reviewed September 18, 2026

Single-function coding challenges such as HumanEval or MBPP do not represent the full difficulty of production software work. Coding agents operate on repository-scale, multi-file codebases and use tools over multiple turns. This guide explains SWE-bench Verified and Terminal-Bench while emphasizing dataset, harness, and contamination limits.

1. The Paradigm Shift: Synthetic Snippets vs. Repository Engineering

Early code evaluation benchmarks tested simple algorithmic functions (e.g., reversing a linked list or checking palindrome strings). However, real-world software engineering requires:

2. SWE-bench Verified Architecture & Methodology

SWE-bench Verified consists of 500 carefully curated, human-validated GitHub issues extracted from major open-source Python repositories including django/django, sympy/sympy, scikit-learn/scikit-learn, and pytest-dev/pytest.

SWE-bench Evaluation Loop:

1. Issue Prompt (Problem Description + Repo State at Commit X)
       │
       ▼
2. Agent Loop (Reads files, edits code, executes pytest commands)
       │
       ▼
3. Git Patch (.patch diff generated by Agent)
       │
       ▼
4. Docker Sandbox Isolation (Applies patch to clean environment)
       │
       ▼
5. FAIL_TO_PASS & PASS_TO_PASS Verification Runner
       │
       ▼
6. Final Score (PASS if fail_to_pass passes AND pass_to_pass has no regressions)
      

3. Why This Page Does Not Republish a Static Leaderboard

Scores change with the agent harness, model snapshot, retry and compute budget, dataset correction, and submission rules. A table copied into an article quickly becomes stale and can compare incompatible setups. Use the official leaderboard for current submissions, then open each result's methodology before drawing a conclusion.

Evidence field Question to ask Common mistake Required record
ModelWhich immutable snapshot was used?Quoting a moving aliasProvider, ID, date
HarnessWhich tools, prompt, retries, and budget?Attributing system score to model aloneCode, config, traces
DatasetWhich revision and exclusions?Mixing Verified, Lite, and other variantsDataset hash
ScorePass@1, best-of-N, or selected run?Ignoring variance and failed setupRuns and denominator

4. Terminal-Bench v2: Evaluating CLI Mechanics

While SWE-bench focuses on Python repository code editing, Terminal-Bench v2 tests an agent's ability to operate inside interactive Linux bash environments. Tasks include:

5. Common Failures & Mitigation Strategies

When running agents on SWE-bench, common failure patterns include:

  1. Context Overflow / Infinite Loops: Agent reads hundreds of files without producing edits. Fixed via AST chunking and search indexing.
  2. Regression Bugs: Agent fixes the target bug but breaks existing functionality. Fixed by enforcing a pre-commit pytest verification run before submitting patches.
  3. Environment Pollution: Modifying global Python packages inside Docker. Fixed via isolated virtualenvs or ephemeral container resets.

6. Build an Internal Coding-Agent Evaluation

  1. Select representative historical tasks you are permitted to use and pin each repository state.
  2. Write executable acceptance tests plus explicit forbidden changes.
  3. Run every candidate in a fresh isolated environment with fixed network and resource policy.
  4. Record tool trajectory, patch, test output, model and harness versions, tokens, cost, latency, and stop reason.
  5. Repeat enough runs to expose variance and inspect every safety or regression failure.
  6. Keep a private holdout and add redacted production failures as regression fixtures.

Measure resolved tasks, regression-free tasks, unauthorized action rate, unnecessary changes, setup failure rate, cost per accepted patch, and time to accepted patch. A patch that passes target tests but deletes unrelated coverage or weakens security is not a success.

7. Primary References

8. Interactive Tools & Benchmarking Resources

To compare live model costs and benchmark scores, explore our interactive utilities:

What a full benchmark run costs

SWE-bench Verified is 500 tasks, and the reason results are often reported from partial runs is that a full run is not cheap. It is, however, cheap enough to be worth knowing the number. Below: one full pass at 60K input, 30K cached and 8K output per task. Per task: 30,000 input tokens at full rate, 30,000 at the cache-read rate, and 8,000 output tokens. One full pass is 500 tasks.

Model Cost per task Full 500-task pass
DeepSeek V4 Pro$0.036$18.15
Claude Sonnet 5$0.146$73.00
GPT-5.6 Sol$0.292$146
Claude Opus 5$0.365$182

A complete run on the mid-tier models is in the low hundreds of dollars, which is within reach of a small team and far below the cost of the engineering time that a misleading partial result can waste. Partial runs are usually a harness limitation, not a budget one.

Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.