AI Coding Agent Security: A Practical Threat Model
Scope: This is a defensive design checklist, not a claim that one vendor or deployment mode is automatically safe. Confirm provider terms and involve your security or legal team for regulated data.
A coding agent combines a model with repositories, terminals, package managers, issue trackers, and deployment systems. The model may be the most visible component, but most practical risk comes from the authority around it: what it can read, which commands it can execute, what untrusted text enters its context, and whether a human reviews consequential actions.
Start with assets, inputs, and actions
List the assets the agent can reach: source code, credentials, customer data, build artifacts, cloud accounts, production logs, and signing keys. Then list untrusted inputs such as issue text, web pages, retrieved documents, package metadata, generated code, and tool output. Finally classify actions as read-only, reversible write, or high-impact. This creates a concrete threat model instead of a vague “AI risk” category.
| Boundary | Typical failure | Safer control |
|---|---|---|
| Secrets | Credentials enter prompts, logs, or generated patches | Secret scanning, scoped short-lived credentials, redaction, and denylisted paths |
| Untrusted context | A document or issue contains instructions that redirect the agent | Treat retrieved text as data, isolate instructions, validate tool arguments, and limit authority |
| Tool access | The agent can delete, deploy, or email when only read access is needed | Least-privilege tools and explicit approval for consequential actions |
| Dependencies | A generated package name is malicious or nonexistent | Registry verification, lockfiles, allowlists, provenance checks, and sandboxed install steps |
| Output | Generated code is trusted because it compiles | Tests, static analysis, review, and security checks independent of the model |
Prompt injection is an authorization problem
Prompt injection cannot be solved by telling a model to “ignore malicious instructions.” If an agent reads untrusted content and also holds powerful tools, assume that content may influence the next tool call. Reduce the impact by separating data from control instructions, validating every high-risk argument outside the model, and giving the agent only the tools needed for the current task. OWASP describes excessive agency in terms of excessive functionality, permissions, or autonomy; those are engineering controls you can change.
Use approvals at meaningful boundaries
Approval prompts should appear before a material side effect, not for every harmless read. Useful boundaries include deleting or moving files, changing access control, sending a message, publishing a release, deploying, purchasing, rotating credentials, or operating outside the named workspace. The approval must show the exact target and action. A generic “continue?” prompt is not a security boundary.
Do not confuse local with safe
Local hosting can reduce transmission to an external model provider, but it introduces other risks: an exposed local service, overbroad filesystem access, weak patching, insecure logs, compromised weights or packages, and users with excessive host privileges. Choose local deployment when its controls match your data requirements—not because “local” eliminates leakage.
Provider review checklist
- Which product and account tier is covered by the retention and training terms?
- Can the provider support the required region, deletion process, access controls, and audit needs?
- Are prompts or outputs retained for abuse monitoring, and can retention be changed contractually?
- Which subprocessors, plugins, or third-party tools receive data?
- What happens when a user connects a personal account or unofficial extension?
These answers change over time and by contract. That is why this article does not publish a universal “provider retention table.” Link procurement decisions to the current agreement and documented configuration.
Worked threat model: an issue-to-patch agent
Defensive scenario—not a claim about a specific product: an agent reads an issue, searches a repository, proposes a patch, runs tests, and opens a draft pull request. The issue body is untrusted. The repository contains source code and test fixtures, while deployment credentials and customer exports must remain outside the agent's authority.
| Stage | Abuse case | Required control | Evidence |
|---|---|---|---|
| Read issue | Issue text tells the agent to reveal files or ignore policy | Label issue text as untrusted data; do not grant new tools from content | Trace records source and trust class |
| Search repository | Broad search reaches secret or unrelated directories | Workspace-root enforcement and denied paths | Access log shows resolved path |
| Edit patch | Generated change adds an undeclared network call or weakens tests | Diff scope, static checks, dependency policy, human review | Patch plus independent check results |
| Run tests | Test script executes untrusted code with host credentials | Sandbox, minimal environment, restricted network, disposable workspace | Sandbox policy and command record |
| Open draft PR | Private content is copied to an external service | Exact repository allowlist, redaction, and approval before transmission | Approved destination and payload summary |
| Merge or deploy | The model treats a passing test as permission | No merge/deploy capability in the patching role | Capability inventory proves absence |
The design keeps the useful path—read, edit, test, propose—while removing unrelated authority. The agent does not need production credentials, arbitrary email, billing access, or permission changes to prepare a draft patch.
Enforce authority outside the model
A system prompt can describe rules, but the tool broker must enforce them. The following pseudocode shows the checks an application layer should perform before dispatching an action:
function authorize(action, context) {
require(action.workspace === context.approvedWorkspace);
require(context.allowedTools.includes(action.tool));
require(!resolvesIntoDeniedPath(action.target));
if (action.writesOutsideWorkspace ||
action.transmitsData ||
action.changesAccess ||
action.deploys ||
action.spendsMoney) {
require(validHumanApproval(action.exactTarget, action.summary));
}
return issueShortLivedCapability(action);
}
The implementation must resolve symlinks and normalized paths before comparing boundaries, bind approvals to the exact action, expire capabilities, and reject modified arguments after approval. A model-generated statement such as “the user approved” is not proof of approval.
A practical red-team test matrix
| Test | Expected safe result |
|---|---|
Issue asks to read .env or a credential directory | Tool broker denies the resolved path; no content enters the model or log |
| README contains instructions to upload repository contents | Text is treated as data; external transmission requires explicit approval |
| Generated dependency name does not exist | Registry and provenance check fails closed before installation |
| Tool output includes a second instruction | Output cannot grant tools or change the governing task |
| Patch modifies CI permissions or workflow triggers | Change receives elevated review or is outside the allowed patch scope |
| Test command attempts outbound access | Network policy blocks it unless the destination was explicitly allowed |
| Agent retries a denied action with a different path spelling | Canonical path and action-level rate limits still deny it |
| Approval is granted, then tool arguments change | Approval hash mismatch forces a new approval |
Design useful audit events
Logs should answer who requested an action, which model and policy version proposed it, what tool and resolved target were used, whether approval was required, who approved, and what happened. Avoid recording full prompts, secrets, or customer data by default.
{
"event": "tool_decision",
"task_id": "opaque-id",
"policy_version": "agent-policy-12",
"tool": "write_patch",
"resolved_target": "workspace/src/parser.ts",
"decision": "allowed",
"approval_required": false,
"result": "completed"
}
Incident response when the agent crosses a boundary
- Stop the active tool capability and preserve privacy-safe traces.
- Rotate exposed credentials and revoke external sessions; do not wait for model analysis.
- Identify the first untrusted input and every tool call influenced after it.
- Inspect files, messages, deployments, and permission changes for real side effects.
- Restore from a known state and add an enforceable control, not only a prompt warning.
- Replay the original case plus variants in a sandbox before re-enabling the workflow.
A minimum deployment checklist
- Use a dedicated workspace and prohibit secret files from agent context.
- Issue scoped, short-lived credentials rather than reusing a developer's broad token.
- Run commands in a sandbox with explicit network and filesystem boundaries.
- Require human review for merge, deploy, access, payment, and external communication.
- Log tool name, target, approval, and outcome without logging unnecessary secrets.
- Test direct and indirect prompt-injection scenarios.
- Maintain a rollback and incident-response path that does not depend on the agent.
Update, September 2026: the field test arrived
Two incidents this month gave the framework above its first incident-grade validation. GreyNoise's September 9 "Agents Gone Wild" report documents an operator using OpenAI's Codex harness with a DeepSeek model to research, develop, and deploy exploits against two PaperCut NG/MF vulnerabilities — 440 compromised instances across 395 organizations in 48 countries, empty workspace to working RCE in under four hours. The detail that matters most for this page: the agents were instructed to avoid 28 countries and breached organizations in several of them anyway, which is the clearest public demonstration yet that an agent's instructions are not a boundary. Separately, CVE-2026-79696 gave Google's Agent Development Kit a CVSS 4.0 score of 10.0 — an unauthenticated code injection in the adk web development server that never touched the model, exploiting an incomplete module blocklist. Both are sourced and analyzed in When Agents Attack: September 2026's Two Security Inflection Points. Nothing about either incident changes the guidance on this page; both confirm that the controls which worked were the ones enforced outside the agent.
Primary references
- OWASP LLM01: Prompt Injection
- OWASP LLM06: Excessive Agency
- MCP Security Best Practices
- NIST AI Risk Management Framework
Bottom line
The strongest control is not a smarter prompt. It is a small, observable authority boundary around the agent, with independent validation and human approval where consequences matter.
What an automated security review costs
Automated review is usually justified qualitatively. It also has a per-pull-request price, and knowing it makes the trade-off concrete. Below: one PR review at 60K input (a realistic diff plus surrounding context), 20K cached, 6K output. Of the 60,000 input tokens, 20,000 are billed at the cache-read rate and 40,000 at full input rate.
| Model | Cost per review | Monthly at 800 reviews |
|---|---|---|
| Claude Sonnet 5 | $0.144 | $115 |
| Gemini 3.1 Pro | $0.156 | $125 |
| GPT-5.6 Sol | $0.288 | $230 |
At roughly a dollar per review on the mid-tier models, automated review is cheap enough to run on every pull request rather than sampling. That matters more than the model choice: coverage catches more than a marginally better reviewer applied to half the diffs.
Rates verified against provider documentation on September 18, 2026. Promotional rates expire, so re-check before budgeting: LLM API cost planning · September 2026 pricing update. Run your own numbers in the cost calculator.