How to Benchmark and Evaluate AI Agents: SWE-bench and Beyond
Evaluating standard language models has evolved from multiple-choice tests (MMLU) and single-function generation (HumanEval) to expert-level reasoning benchmarks like Humanity's Last Exam (HLE). However, evaluating an AI Agent is far more complex. Agents operate in loops, write multi-file patches, invoke shell commands, and read local errors. Evaluating their terminal interaction capability requires benchmarks like Terminal-Bench v2. This article explains how the AI community benchmarks agentic systems using SWE-bench and Terminal-Bench v2, and how you can set up evaluation rigs for your own agents.
SWE-bench: The Ultimate Coding Test
A widely used benchmark for evaluating AI coding agents is SWE-bench. Instead of isolated algorithm prompts, it builds tasks from real issue and repository histories in open-source Python projects.
The agent is given a codebase and an issue description. It must navigate the directory, edit the code, and output a patch. The evaluation harness runs the project's existing test suite plus new tests specifically written for that issue. The agent is marked successful only if the test suite compiles and passes.
How to Read an Agent Benchmark Result
| Field | Why it changes the score |
|---|---|
| Dataset and revision | Tasks can be corrected, removed, public, or contaminated |
| Model snapshot | An alias may move after the evaluation date |
| Agent harness | Search, tools, prompts, memory, and retry policy affect outcomes |
| Compute budget | More turns, samples, or retries can raise pass rate and cost |
| Environment | Repository setup, network, dependencies, and tests can fail independently |
| Scoring | Pass@1, best-of-N, selected runs, and averages answer different questions |
Three Pillars of Agent Evaluation
If you are building custom AI agents for your team or organization, relying on generic benchmarks isn't enough. You need an internal testing harness composed of three layers:
1. Secure Sandboxed Execution
Never run agent-generated commands or scripts directly on your host machine. Safe evaluation requires containerized environments (like Docker containers or microVMs) that isolate file modifications and command execution. After each run, discard the environment and restore it from a versioned clean image; verify the reset procedure rather than assuming isolation alone removed every artifact.
2. Rubric-Based Grading
Some outputs, such as design quality or documentation usefulness, require judgment. Use blinded human labels for the highest-impact cases. A model grader can help scale a detailed rubric, but calibrate it against humans, require evidence, randomize candidate order, and permit an uncertain result. Never let a model-graded style score override failing tests, an authorization violation, or an unsafe side effect.
3. Golden Datasets
Create a static list of 20 to 100 internal coding tasks that represent the typical workload of your team. This dataset should include:
- The initial state of the codebase.
- The developer instruction.
- The expected modified files.
- Assertion scripts or unit tests to validate the change.
Run your agent against this dataset after making modifications to the agent's system prompt or toolset to detect performance regressions.
Evaluate the Complete Trajectory
Store the initial state, every model turn, tool proposal, validated argument, tool result, approval, state mutation, final output, token use, latency, and stop reason. A final patch can pass while the trajectory reads secrets, reaches an unapproved host, or spends far beyond budget. Score both outcome and process constraints.
Include tool timeouts, stale repository revisions, prompt injection inside issues, dependency failures, duplicate job delivery, budget exhaustion, and cancellation. Every run should terminate as completed, declined, needs input, needs approval, or failed—never an invisible loop.
Measuring the Cost-to-Performance Ratio
When comparing agents, code correctness is only part of the equation. You must also track run cost and time-to-resolution. An agent that solves 50% of tasks but spends $10 in API fees per run may be less useful than a faster agent that solves 40% of tasks at a cost of $0.10 per run.
Primary References
- Anthropic: Demystifying evals for AI agents
- SWE-bench official documentation and leaderboards
- Terminal-Bench official site
The Verdict
Building high-performance AI agents requires treating them like human software engineers. Do not evaluate them based on static test templates. Instead, build a sandboxed testing harness that runs unit tests, tracks token costs, and runs regular regression checks against your team's real-world code workloads.
What it costs to run an evaluation suite
Evaluation is the one agent activity where cost is easy to underestimate, because a full suite re-runs the same harness hundreds of times and each run re-sends context. Budget for the suite, not for the single task. Below: one full suite run at 200K input tokens, 120K of them cached, and 20K output. Of the 200,000 input tokens, 120,000 are billed at the cache-read rate and 80,000 at full input rate.
| Model | Cost per suite run | Monthly at 40 suite runs |
|---|---|---|
| DeepSeek V4 Pro | $0.095 | $3.80 |
| Claude Sonnet 5 | $0.384 | $15.36 |
| Gemini 3.1 Pro | $0.424 | $16.96 |
| GPT-5.6 Sol | $0.768 | $30.72 |
Forty suite runs a month is a realistic cadence for a team that evaluates on every release, and it costs less than most teams spend on CI compute. The expensive failure mode is not the suite itself but re-running it without a cache, which roughly triples the input line.
Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.