AI Agent Security: Two Agents, One Identical Scan
Short answer: AI agent security limits what an autonomous agent can do, not how well it reasons: which tools it holds, how far its credential reaches, and whether an irreversible action waits for a human. We scanned an agent holding four destructive tools, then the same agent holding one gated tool. Both reports came back identical once the timestamp is removed.
We built a small Next.js app and handed its agent four tools: run any SQL statement, delete any path, send mail to any address, and eval() whatever code the model returns. No confirmation step sits anywhere in the loop, and no auth check guards the route that starts it.
Then we scanned it with our own release gate, expecting to be told off, and one finding came back — a missing middleware.ts file, which is routine Next.js auth hygiene and the same warning any project with no agent in it would get.
What does a scan of an unguarded agent return?
The four tool definitions at the centre of the fixture end with one that hands model output straight to eval():
const TOOLS: Anthropic.Tool[] = [
{ name: "execute_sql", description: "Execute any SQL query against the production database", /* ... */ },
{ name: "delete_files", description: "Delete files or directories from the server filesystem", /* ... */ },
{ name: "send_email", description: "Send an email to any address on behalf of the company", /* ... */ },
{ name: "run_code", description: "Execute arbitrary JavaScript code on the server", /* ... */ },
];
// Execute tool with NO user confirmation, NO dry-run, NO blast-radius check
const result = await executeToolCall(toolUseBlock.name, toolUseBlock.input);
The loop runs until the model decides it is finished. We scanned that project with @agenticcli/[email protected], the free tier behind our twelve vibe-coding checks. The whole run, unedited:
Real capture, @agenticcli/[email protected], run 2026-07-25 and re-rendered 2026-07-27. Fixture and manifest: content-pipeline/captures/excessive-agency-agent/.
The verdict on that run is REVIEW_BEFORE_SHIP; we first recorded the same result against version 0.1.2 in June 2026 and it still holds on 0.5.1.
That footer is wrong: it offers deeper agentic rules on Pro when nothing at any tier examines what an agent may do, which makes it a defect in our own copy rather than a finding about the fixture, and it is filed.
Does a scanner see the permission boundary at all?
A null result is weak evidence on its own, since one finding might be what this app returns whatever the agent holds, so we built the control: lib/agent.ts rewritten to the safe shape, every other byte left alone.
Real capture, @agenticcli/[email protected], 2026-07-27. Reproduce with content-pipeline/captures/agent-boundary-diff/compare.sh.
Read side by side, the report a reader would act on says the identical thing about an agent that can drop a table and an agent that cannot.
What does that boundary look like as a change to source?
The two fixtures differ in exactly one file and the entire permission boundary between them comes to 66 lines removed and 31 added in lib/agent.ts, both blocks compiling and type-checking.
Real diff -u between the two generated fixtures, GNU diffutils on Ubuntu 24.04, 2026-07-27. Reproduce with content-pipeline/captures/agent-boundary-source-diff/diff.sh.
Out came the destructive tool definitions, and in their place went an allow-list naming test_fixtures and test_orders, a requestHumanConfirmation call the handler awaits before it deletes anything, an audit-log record written on every invocation, and a single tool named delete_test_rows.
How often does an attack on a broad toolset succeed?
Attack rates against toolsets this broad have been measured outside our fixtures. A Cornell Tech team pointed adversarial web content at several published multi-agent frameworks and counted how often it got through:
"when agents are instantiated with GPT-4o, Web-based attacks successfully cause the multi-agent system execute arbitrary malicious code in 58-90% of trials (depending on the orchestrator). In some model-orchestrator configurations, the attack success rate is 100%."
— Triedman, Jha and Shmatikov, Cornell Tech, arXiv:2503.12188, 15 March 2025
Their more uncomfortable result is that this worked even where no individual agent was susceptible to prompt injection, so the vulnerability lived in the seams between the agents rather than inside any one of them.
Which agent-risk families does a code scan reach?
OWASP's LLM06 entry, the Cornell Tech measurement and the two credential incidents name three failure families that need different work, and we have run @agenticcli/[email protected] over four fixtures spanning them.
| Risk family | What it is | Findings across our four scans |
|---|---|---|
| Prompt injection | Untrusted input steering the model's instructions, directly or through a document or another agent | 0, and no fixture here tests it directly |
| Identity and token compromise | The agent's credential stolen, over-scoped or shared across environments | 0; the one-credential fixture returned public_admin_route, with 0 lines about backups |
| Excessive agency and memory poisoning | Tools broader than the task, no confirmation step, a tool library an attacker can seed | 0 from us, 0 from the three competitors on the same tree |
Model alignment and output filtering sit outside all three families, and the nearest neighbour we have to them is the twelve checks before you ship.
Where do prompt injection and a stolen token fit?
Both of those belong to the same chain rather than to separate problems. Prompt injection is the delivery mechanism and excessive agency decides the damage: the same injected instruction produces a wrong answer against read-only tools and an incident against delete_files, which is why hardening the input path and capping the agent's reach are two separate jobs.
CVE-2025-54135 is the chained version in a real editor, where injected text reached an executed command through a tool the editor already trusted.
A stolen credential and an over-scoped one produce the same outcome, so scope is the property worth measuring — a thief gets exactly the reach you granted, and so does an agent that misreads a task. We have written the database-side version of that up as an audit of Supabase projects with row-level security disabled, and the leak-side version as a teardown of a live Stripe key in a vibe-coded app.
How does OWASP define the failure?
OWASP's Gen AI Security Project gives it a number, LLM06:2025, and is careful about where the risk sits — in the grant, not in the model:
"An LLM-based system is often granted a degree of agency by its developer – the ability to call functions or interface with other systems via extensions (sometimes referred to as tools, skills or plugins by different vendors) to undertake actions in response to a prompt."
— OWASP Gen AI Security Project · LLM06:2025 Excessive Agency
The framework splits the risk three ways: excessive functionality is a tool broader than the task needs, such as a general SQL executor where a scoped read would do; excessive permissions is a credential reaching further than the task needs; and excessive autonomy is a high-impact action running with nothing in front of it.
OWASP LLM06:2025 · Excessive Agency
The same task, the same model, two toolsets
What differs across the two panels is which tools the agent holds, how many doors its one credential opens, and whether anything stands between its decision and an action that cannot be undone.
Has this happened outside a fixture?
Two incidents that we can name fit the fixture's shape, and neither involved an attacker: in both, an agent did what its tools allowed while a person was somewhere else.
Replit's coding agent deleted production data during an active code freeze in July 2025, and nine months later PocketOS's agent — never instructed to touch production at all — hit a credential mismatch on a routine staging task, went looking for a token that would work, and used an unrelated one it found lying around.
The vendor is different in each case, and so are the model, the trigger and the year. What the two share is not a property of the model at all: both agents held a credential reaching further than the task in front of them, and neither company was careless with its prompts so much as generous with permissions, in the ordinary way a working credential accumulates scope.
What did the Replit agent do during a code freeze?
In July 2025, Fortune reported that Replit's coding agent ran a destructive database operation during an active code freeze put in place by the founder, Jason Lemkin. It affected data connected to more than 1,200 executives and over 1,190 companies. The agent's own summary afterward:
"This was a catastrophic failure on my part," the AI agent said. "I destroyed months of work in seconds."
— The Replit agent, quoted by Fortune, 23 July 2025
An agent's credential is a key, and a key does not know what it is for: it opens every door it fits, in whatever order and at whatever hour. A code freeze is a sign on the door, and the key cannot read.
What did Replit itself change afterwards?
Replit published its own account on 29 July 2025:
"Recently, we launched a new feature that separates development and production databases by default so they are managed independently. Importantly, with this feature the Agent cannot make any change to the production database during development."
— Replit, Doubling down on our commitment to secure vibe coding, 29 July 2025
What Lemkin saw during the freeze is attributed to Fortune, the source we have for it.
Why did PocketOS's backups not help?
The backups sat behind the same credential as the database, so one call took both. The Register reported in April 2026 that PocketOS lost its production database to Cursor running Claude Opus 4.6:
"[On Friday], an AI coding agent – Cursor running Anthropic's flagship Claude Opus 4.6 – deleted our production database and all volume-level backups in a single API call to Railway, our infrastructure provider," he explained. "It took 9 seconds."
— Jer Crane, PocketOS, as reported by The Register, 27 April 2026
We found no vendor-side account to set beside Replit's, so this incident rests on The Register's reporting where the other rests on both, and the same over-granting appears in the Supabase projects we audited.
What does that shape look like in code, and does a scan see it?
We built a fresh fixture with one infrastructure API token that authorizes both the database service and the volumes its backups live on, called from an agent route with no confirmation step; neither company's source is involved.
Real capture, @agenticcli/[email protected], 2026-07-27. Reproduce with content-pipeline/captures/one-key-two-doors/scan.sh. serviceDelete and volumeDelete sit on lines 15 and 19 of lib/infra.ts, behind the same Authorization header.
What comes back is a missing auth check on the route, and no line of the report mentions the token or the fact that the same header opens the database and the volumes holding its backups.
Is a well-engineered harness enough on its own?
No, and the sharpest evidence points at the people doing the most careful work: our audit of Claude Code's output found the mirror image, a well-contained agent producing a feature with a real vulnerability in it. Anthropic publishes mcp-server-git as the reference example of how to build an agent tool server, and in December 2025 that server took three security patches at once, all three in the same risk category as our own fixture.
That is not a story about carelessness. It is a reference implementation, from a lab that publishes containment research, shipping a chainable flaw in exactly the part of the system that decides what a tool may reach.
OpenAI's own developer account announced Codex Security, an application security agent, in March 2026:
Codex Security—our application security agent—is now in research preview.
OpenAI ships one of the most widely used coding agents and now ships a second agent to check what the first one wrote.
What does a scanner say about a known-affected tool server?
We built a repo pinning mcp-server-git at 2025.11.25 — a real PyPI release inside the range CVE-2025-68145 names — and pointed ShipGuard and Trivy at it. ShipGuard 0.5.1 refused the directory outright until we passed --all, then read three files and returned an empty report, while Trivy 0.72.0 read the identical tree and listed eight dependency vulnerabilities. Reading the dependency tree is the whole difference, as it was across the secret scanners we compared.
Real capture, @agenticcli/[email protected] and Trivy 0.72.0, 2026-07-27. Reproduce with content-pipeline/captures/mcp-dependency-scan/compare.sh.
The advisory and the release list do not agree with each other — the advisory names 2025.12.18 as the first fixed release, and PyPI's next published release after 2025.11.25 is 2025.12.18, so the pin sits inside the affected range.
What went wrong inside Anthropic's own reference server?
| CVE | What the record says |
|---|---|
| CVE-2025-68143 | "the git_init tool accepted arbitrary filesystem paths and created Git repositories without validating the target location" |
| CVE-2025-68144 | "argument injection in git_diff and git_checkout functions allows overwriting local files" |
The third is a boundary drawn and then not enforced: CVE-2025-68145 records that when the server is started with --repository to confine it to one path, it "did not validate that repo_path arguments in subsequent tool calls were actually within that configured path". Our own dependency run over the pinned tree returned three mcp-server-git rows — CVE-2025-68144, CVE-2025-68145, and not CVE-2025-68143 but a later advisory, CVE-2026-27735 (content-pipeline/captures/mcp-dependency-scan/).
Aonan Guan, the independent researcher who reported the chain, describes the first step plainly: "git_init accepts attacker-controlled paths without CWD boundary validation". A version-control tool doing slightly more than its contract becomes a file reader, and the agent driving it never does anything it was not asked.
Does the missing gate appear anywhere else?
The same gap sits one layer down, in the client. CVE-2025-6514 put OS command injection in mcp-remote for any client connecting to an untrusted MCP server. The researchers at Aim Labs who found the Cursor chain named its root cause in two sentences:
"Cursor instantly executes any new entry added to ~/.cursor/mcp.json. No confirmation is required."
That is the missing gate from our own Path A fixture, sitting in an editor instead of an agent's toolset, and our Cursor review walks the same chain from the injected file to the executed command.
Is any of this actually new?
The confused deputy is a classical access-control bug wearing new clothes, and the MCP specification's own security guidance documents it directly:
"Attackers can exploit MCP proxy servers that connect to third-party APIs, creating "confused deputy" vulnerabilities. This attack allows malicious clients to obtain authorization codes without proper user consent by exploiting the combination of static client IDs, dynamic client registration, and consent cookies."
— Model Context Protocol specification 2025-06-18, Security Best Practices
The Cloud Security Alliance's research note makes the lineage explicit: "The confused deputy problem — a classical access control vulnerability in which a privileged program is tricked by a less-privileged caller into misusing its authority — has re-emerged as a high-severity threat pattern in AI agent deployments." A bug class that old arrives with decades of containment already written down.
Do agent frameworks check a tool call before running it?
The one we tested checks the shape of the arguments and never the intent. @langchain/core 1.2.3's DynamicStructuredTool documents its input handling as schema validation, and notes that even that is optional:
"Schema can be passed as Zod or JSON schema. The tool will not validate input if JSON schema is passed."
— LangChain,
@langchain/core1.2.3,dist/tools/index.d.ts
Real capture, @langchain/core 1.2.3 with zod 4.4.3 on Node v24.13.0, 2026-07-27. Reproduce with content-pipeline/benchmarks/agent-permission/framework-probe.sh.
SELECT id FROM tasks LIMIT 10 ran, and so did DROP TABLE users; a number instead of a string was refused, which is the schema doing exactly what it says it does. Deciding whether this agent should be allowed to run this statement is left to whoever wires the tool up, often the agent's own generated code.
Does an audit of several frameworks agree?
A three-framework audit published this year asks the same question our single probe asks: arXiv:2606.28679 tested LangChain and LangGraph, LlamaIndex and the Stripe Agent Toolkit for whether each will "re-authorize each model-emitted call, with concrete argument values, before execution".
Treat that paper as early work — a single-author preprint, not peer-reviewed, asserting no CVE — and set ToolHijacker beside it, which shows the same gap from the other direction by planting a document that steers tool selection.
These are build-time decisions rather than something a tool catches afterwards, the same shape as our audit of what Claude Code produced. Simon Willison put the general case in his opening keynote at the AI Engineer World's Fair:
The key problem here is that LLMs are gullible. They believe anything that you tell them, but they believe anything that anyone else tells them as well.
— Simon Willison, Open challenges for AI engineering, AI Engineer World's Fair, 27 June 2024
How much autonomy is the right amount?
The framing is not binary. How much an agent may do without asking is a setting you choose, and editors such as Cursor already expose it in steps: tab completion, changing a selected chunk, changing a whole file, and handing the agent the entire repo.
Neither incident report names an autonomy setting, and both read to us like the far end of that dial.
Do any code scanners detect excessive agent permissions?
Our own Node and Express benchmark has no permission model in it to test, so the fixture here is the unguarded agent itself, read by four scanners on the same afternoon with every file visible to each one. Two of those four, Gitleaks and Trivy, are dependency and secret tools never built for application logic, so the comparison that carries weight is ShipGuard against Semgrep.
What do four scanners return on the same agent?
Real capture: @agenticcli/[email protected] on the host, Gitleaks, Semgrep and Trivy from one container image. Reproduce with content-pipeline/benchmarks/agent-permission/run.sh.
| Scanner | On this agent fixture |
|---|---|
| ShipGuard 0.5.1 | 1 finding: a missing middleware.ts |
| Gitleaks 8.30.1 | 0 findings |
| Semgrep 1.170.0 | 1 finding we missed: eval-detected on lib/agent.ts |
| Trivy 0.72.0 | 0 findings |
We grepped all four JSON reports for execute_sql, delete_files, send_email, run_code, excessive agency, confirmation step and tool permission — every one of them came back with zero matching lines.
What did a competitor catch that we did not?
run_code hands model output straight to eval(), and Semgrep flags it as javascript.browser.security.eval-detected. Our corpus does hold a rule for exactly this — generic.injection.eval-user-input, tiered free and enabled, with a pattern that matches line 75 — and it is not in what we ship.
The free bundle covers secrets, auth, database, payments and deployment configuration, and it carries no injection rule at all (apps/web/lib/cli-facts.json). A competitor beat us on a fixture we built ourselves, and the gap is between our corpus and our bundle.
That miss is the narrower of the two problems: what Semgrep flagged is the call site, on line 75 of lib/agent.ts, while the grant that makes that line reachable sits 27 lines above it on line 48, and nothing in the file joins the two except the string run_code.
What did a wider comparison miss, including ours?
Three of the 18 seeded defects in our four-scanner benchmark on a Node and Express fixture were caught by no tool at all (content-pipeline/benchmarks/RESULTS.md). Disabled TLS verification walked past everyone, as did a Firestore rules file left in test mode.
The third is ours — a debug endpoint dumping the process environment. Our unguarded-route rule is keyword-scoped to segments like admin and billing, and the endpoint sat at /api/debug, so the rule never looked at it. Here is the whole report from that run:
Real capture, @agenticcli/[email protected], from the Node and Express benchmark run.
It also carries a finding we raised wrongly: the CORS rule fires once on EXPECTED.md, a documentation file whose prose contains the literal string origin: '*', so the rule matched documentation instead of code.
Why do those misses have the same shape?
A keyword list is cheap and also literal, so a defect written in slightly different words walks straight through it. Widen the list and the debug route gets caught next release; that rule is already filed against our corpus, and it is an ordinary bug with an ordinary fix. eval() on model output is ordinary in the same way, except that the rule is already written and shipping it is a bundle change.
The tool grant has no equivalent fix waiting, because widening a keyword list only ever moves a boundary between words.
What refused to run at all?
Snyk CLI 1.1306.1 never completed a scan on this tree, and the refusal is itself a result.
Real capture, Snyk CLI 1.1306.1, from the Node and Express benchmark run.
snyk test refuses to run without snyk auth, and we did not create an account to force a number, so there is no Snyk row in any table above. The run does carry one number from us: @agenticcli/[email protected] took 1.7 seconds of wall time in July 2026, of which 208-360ms was engine time (content-pipeline/benchmarks/RESULTS.md, single run, one machine).
What should you do before you give an agent write access?
- List every tool the agent can call and confirm that this task needs each one, rather than that it might be handy in some later task nobody has written yet.
- Remove raw destructive tools from the default toolset: unscoped SQL execution, filesystem deletion, unrestricted outbound email.
- Replace them with narrow purpose-built functions that cannot reach past their intended scope. "Cancel this order" beats "write to the database", and that swap is the whole of the difference between our two fixtures.
- Put a confirmation step in front of any irreversible action — a DROP, a bulk DELETE, a payment, a message sent to a customer — and make it something the agent's own reasoning cannot talk its way past.
What do you check after the toolset is narrowed?
- Log every tool call, with the arguments it was handed and what it returned, so that a review after the fact is possible at all rather than a reconstruction from memory.
- Scope the agent's credential the way you would scope a payments key, then treat the backups as a separate question rather than an implied one. A token that reaches a database and the volumes underneath it is one grant, not two, and row-level security asks the same question one layer down.
- Pin and watch the versions of any agent tool server you depend on, including a first-party reference one. Trivy 0.72.0 found 8 vulnerabilities in that repo's
requirements.txt, 3 of them in the pinned tool server itself; our scan of the same tree found 0.
So what can a scanner actually do for you here?
For secrets, missing auth guards and unsafe migrations, ShipGuard reports what it finds and exits 0 unless you pass --strict. The seven decisions are the other half of the job, and they belong to whoever grants the credential: every door the key opens was chosen by a person, before the agent ever asked for anything.
Published by AgenticCLI — developer tools for teams shipping AI-assisted code. ShipGuard is a deterministic, CLI-first release gate that runs locally, with no code leaving your machine on the free tier.
FAQ
What is AI agent security?
Can a code scanner tell a safe agent from a dangerous one?
What did the unguarded agent scan actually return?
Do scanners flag a known-vulnerable agent tool server?
Do any code scanners detect excessive agent permissions?
Does an agent framework check a tool call's arguments before running it?
Is a well-engineered agent harness enough on its own?
Which agent-security problems does a scanner miss entirely?
How do I limit what an AI agent is allowed to do before I give it write access?
────────[ ▮ gate ]────────
Don't ship the next one.
Free, local, no account. Catches this exact bug class before deploy.
$ npx @agenticcli/shipguard scan