Deep Dive into SWE-bench Verified & Terminal-Bench v2
Single-function coding challenges such as HumanEval or MBPP do not represent the full difficulty of production software work. Coding agents operate on repository-scale, multi-file codebases and use tools over multiple turns. This guide explains SWE-bench Verified and Terminal-Bench while emphasizing dataset, harness, and contamination limits.
1. The Paradigm Shift: Synthetic Snippets vs. Repository Engineering
Early code evaluation benchmarks tested simple algorithmic functions (e.g., reversing a linked list or checking palindrome strings). However, real-world software engineering requires:
- Navigating codebase directory trees containing thousands of files.
- Locating relevant class definitions and function call chains across decoupled modules.
- Modifying multi-file state without breaking downstream dependencies or existing unit tests.
- Executing terminal commands, installing dependencies, and analyzing test runner stack traces.
2. SWE-bench Verified Architecture & Methodology
SWE-bench Verified consists of 500 carefully curated, human-validated GitHub issues extracted from major open-source Python repositories including django/django, sympy/sympy, scikit-learn/scikit-learn, and pytest-dev/pytest.
SWE-bench Evaluation Loop:
1. Issue Prompt (Problem Description + Repo State at Commit X)
│
▼
2. Agent Loop (Reads files, edits code, executes pytest commands)
│
▼
3. Git Patch (.patch diff generated by Agent)
│
▼
4. Docker Sandbox Isolation (Applies patch to clean environment)
│
▼
5. FAIL_TO_PASS & PASS_TO_PASS Verification Runner
│
▼
6. Final Score (PASS if fail_to_pass passes AND pass_to_pass has no regressions)
3. Why This Page Does Not Republish a Static Leaderboard
Scores change with the agent harness, model snapshot, retry and compute budget, dataset correction, and submission rules. A table copied into an article quickly becomes stale and can compare incompatible setups. Use the official leaderboard for current submissions, then open each result's methodology before drawing a conclusion.
| Evidence field | Question to ask | Common mistake | Required record |
|---|---|---|---|
| Model | Which immutable snapshot was used? | Quoting a moving alias | Provider, ID, date |
| Harness | Which tools, prompt, retries, and budget? | Attributing system score to model alone | Code, config, traces |
| Dataset | Which revision and exclusions? | Mixing Verified, Lite, and other variants | Dataset hash |
| Score | Pass@1, best-of-N, or selected run? | Ignoring variance and failed setup | Runs and denominator |
4. Terminal-Bench v2: Evaluating CLI Mechanics
While SWE-bench focuses on Python repository code editing, Terminal-Bench v2 tests an agent's ability to operate inside interactive Linux bash environments. Tasks include:
- Configuring Nginx reverse proxies and SSL certificates.
- Debugging complex C++ build failures and CMake targets.
- Executing multi-step database migrations and SQL schema fixes.
- Handling shell signal handlers, process management, and environment variables.
5. Common Failures & Mitigation Strategies
When running agents on SWE-bench, common failure patterns include:
- Context Overflow / Infinite Loops: Agent reads hundreds of files without producing edits. Fixed via AST chunking and search indexing.
- Regression Bugs: Agent fixes the target bug but breaks existing functionality. Fixed by enforcing a pre-commit
pytestverification run before submitting patches. - Environment Pollution: Modifying global Python packages inside Docker. Fixed via isolated virtualenvs or ephemeral container resets.
6. Build an Internal Coding-Agent Evaluation
- Select representative historical tasks you are permitted to use and pin each repository state.
- Write executable acceptance tests plus explicit forbidden changes.
- Run every candidate in a fresh isolated environment with fixed network and resource policy.
- Record tool trajectory, patch, test output, model and harness versions, tokens, cost, latency, and stop reason.
- Repeat enough runs to expose variance and inspect every safety or regression failure.
- Keep a private holdout and add redacted production failures as regression fixtures.
Measure resolved tasks, regression-free tasks, unauthorized action rate, unnecessary changes, setup failure rate, cost per accepted patch, and time to accepted patch. A patch that passes target tests but deletes unrelated coverage or weakens security is not a success.
7. Primary References
- OpenAI introduction to SWE-bench Verified
- SWE-bench official documentation and leaderboards
- Terminal-Bench official site
- OpenAI audit of SWE-bench Verified limitations
8. Interactive Tools & Benchmarking Resources
To compare live model costs and benchmark scores, explore our interactive utilities:
- 📊 Interactive AI Agent Benchmark Matrix
- 🧮 LLM API & Token Cost Calculator
- 🛠️ AI Agent Builder Lab & Code Studio
What a full benchmark run costs
SWE-bench Verified is 500 tasks, and the reason results are often reported from partial runs is that a full run is not cheap. It is, however, cheap enough to be worth knowing the number. Below: one full pass at 60K input, 30K cached and 8K output per task. Per task: 30,000 input tokens at full rate, 30,000 at the cache-read rate, and 8,000 output tokens. One full pass is 500 tasks.
| Model | Cost per task | Full 500-task pass |
|---|---|---|
| DeepSeek V4 Pro | $0.036 | $18.15 |
| Claude Sonnet 5 | $0.146 | $73.00 |
| GPT-5.6 Sol | $0.292 | $146 |
| Claude Opus 5 | $0.365 | $182 |
A complete run on the mid-tier models is in the low hundreds of dollars, which is within reach of a small team and far below the cost of the engineering time that a misleading partial result can waste. Partial runs are usually a harness limitation, not a budget one.
Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.