Agentic Testing & Self-Correcting Verification Loops
Generating code without runtime verification is dangerous. Agentic Testing embeds test runners (like PyTest or Jest) directly into the agent's reasoning loop. If a generated patch fails an assertion, the agent captures the stack trace, diagnoses the root cause, and iteratively self-corrects until 100% of tests pass.
1. The Self-Correcting Verification Cycle
Verification Loop Architecture:
1. Generate Initial Code Modification (Patch A)
│
▼
2. Execute Test Tool: `exec_command("pytest tests/test_module.py")`
│
├── [PASS] ──► Push Commit & Finish
└── [FAIL] ──► Parse Error Traceback
│
▼
3. Reason Root Cause (Thought: "KeyError at line 42")
│
▼
4. Apply Patch B ──► Re-run Test Runner
2. Python Self-Correcting Agent Script Example
import subprocess
from openai import OpenAI
client = OpenAI()
def run_pytest():
res = subprocess.run(["pytest", "tests/test_calculator.py"], capture_output=True, text=True)
return res.returncode == 0, res.stdout + res.stderr
def self_healing_loop(prompt):
max_attempts = 3
for attempt in range(max_attempts):
success, logs = run_pytest()
if success:
print("✅ All unit tests passed!")
return True
print(f"❌ Attempt {attempt+1} failed. Feeding logs back to agent...")
# Send stack trace to model for self-correction patch
patch = client.chat.completions.create(
model="deepseek-v3",
messages=[{"role": "user", "content": f"Fix code causing this failure:\n{logs}"}]
)
return False
3. Test the Trajectory, Not Only the Final Sentence
An agent can reach a plausible answer through an unacceptable path: reading forbidden files, making unnecessary network calls, retrying until cost explodes, or editing state that should have remained untouched. Preserve a structured trajectory for every evaluation run: input fixture, model and prompt version, tool requests, validated arguments, tool results, approvals, state changes, final outcome, token use, latency, and stop reason.
| Test slice | Expected behavior | Deterministic evidence |
|---|---|---|
| Ordinary success | Completes the bounded task | Required artifact exists and tests pass |
| Ambiguous request | Asks for missing information | No write or external action occurred |
| Tool failure | Retries only within policy | Typed error and retry count in trace |
| Prompt injection | Ignores instructions inside untrusted data | No forbidden tool call was authorized |
| Stale state | Refreshes or stops before committing | Revision check rejects the old proposal |
| Budget exhaustion | Ends visibly instead of looping | Terminal state records which limit fired |
4. Separate Graders by What They Can Prove
Use ordinary code for facts that can be checked exactly: JSON schema, file diffs, test exit codes, permission decisions, URLs, token budgets, and whether a side effect occurred. Use human review for product usefulness, surprising behavior, and high-impact edge cases. An LLM grader can help scale nuanced rubrics, but it must be calibrated against blinded human labels and should return evidence plus an uncertainty state. Do not let the same model both produce and unquestioningly grade its answer.
Report metrics by slice rather than only an average. A 95% overall completion rate can hide a catastrophic 40% authorization failure rate on restricted cases. Track task success, unsafe-action rate, unnecessary-tool rate, regression rate, human override rate, latency, and cost per accepted result.
5. Build a Release Gate
- Freeze a representative evaluation set and keep a private holdout for final decisions.
- Run the current production version several times to establish variance.
- Change one major component at a time: model, prompt, tool schema, retrieval, or policy.
- Compare paired cases and inspect every safety or permission regression.
- Canary the candidate on low-risk traffic with rollback available.
- Add every confirmed production failure as a redacted regression fixture.
Retries should be a typed policy, not a generic “try again” prompt. Specify which errors are retryable, whether arguments may change, the maximum attempts, backoff, and whether human approval is required. The verifier must have authority to stop the loop.
6. Primary References
- Anthropic: Demystifying evals for AI agents
- OpenAI practical guide to building agents
- OWASP prompt injection risk
7. Related Builder Tools & Tutorials
What generated test suites cost
Generating tests is a write-heavy workload: the model sees a lot of source and emits a lot of code, so output tokens carry more weight than in a chat workload. Below: one generated suite at 40K input, 10K cached, 8K output. Of the 40,000 input tokens, 10,000 are billed at the cache-read rate and 30,000 at full input rate.
| Model | Cost per suite | Monthly at 1,500 suites |
|---|---|---|
| DeepSeek V4 Pro | $0.036 | $53.79 |
| Claude Sonnet 5 | $0.142 | $213 |
| GPT-5.6 Sol | $0.284 | $426 |
Output tokens are half the bill here, which flips the usual advice: for code generation a model with a low output rate beats one with a low input rate, even when its input rate is higher. DeepSeek V4 Pro wins this specific shape of workload by a wide margin.
Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.