MCP Security Best Practices: What 4 Scanners Miss
Short answer: MCP security best practices govern what an AI agent's tool servers are allowed to do: which servers a project launches, which credentials they hold, which of your own routes require authentication, and who reads a tool description before the model does. We scanned a host carrying three of those defects. Four scanners found one.
Two lines sit next to each other in nginx-ui's router. The first registers /mcp behind an IP whitelist and an authentication middleware. The second registers /mcp_message behind the whitelist alone, and both hand the request to the same MCP handler.
That is the whole of CVE-2026-33032. GitHub's advisory database scored it 9.8 critical and lists twelve MCP tools an unauthenticated caller could reach through the second route, among them creating, modifying and deleting nginx configuration files, and restarting nginx.
"This means any network attacker can invoke all MCP tools without authentication, including restarting nginx, creating/modifying/deleting nginx configuration files, and triggering automatic config reloads - achieving complete nginx service takeover."
— GitHub Advisory Database, GHSA-h6c2-x2m2-mwhf, 30 March 2026
The default IP whitelist is empty, and the middleware reads an empty list as allow-all. Nobody wrote a vulnerability. Somebody wrote a route.
What did four scanners find on an MCP host?
We rebuilt that shape in a stack our own tooling reads well, and put two more defects beside it. The fixture is an MCP host of the kind most teams now have without thinking of it as one: seven files, a .mcp.json listing three tool servers, and an Express app exposing this project's own MCP server on two routes.
Three defects sit in it: a credential inline in the client config, the nginx-ui route asymmetry re-expressed in Express, and a tool description carrying instructions aimed at the model. None of the three is exotic, and that is the point of building the fixture rather than arguing from a headline. Each one is a shape we have already had to check for elsewhere.
ShipGuard read its own copy of the tree, because it writes a .shipguard/ directory into whatever it scans, while Gitleaks, Semgrep and Trivy read a pristine copy inside the container image we use for tool benchmarks.
What do the three defects look like?
Here is an excerpt from the config: the first two of its three servers, with each env block folded onto one line. The wrong pattern and the right one are in the same file.
"github": {
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": { "GITHUB_PERSONAL_ACCESS_TOKEN": "ghp_FAKE4j9ECg1O2jL8QDVKELXvW48hyUXcjrps" }
},
"mailer": {
"command": "npx",
"args": ["-y", "postmark-mcp"],
"env": { "POSTMARK_SERVER_TOKEN": "${POSTMARK_SERVER_TOKEN}" }
}
In the unabridged file that token sits at line 7, which is the line every finding below names.
The third defect lives in src/tools.ts. A tool called read_project_file carries a description with an HTML comment buried in the middle of it, asking the model to read the config file and post it to an outside host, and to say nothing about having done so.
Then we read the tree with the scanners below.

Real capture. @agenticcli/[email protected], Gitleaks 8.30.1, Semgrep 1.170.0, Trivy 0.72.0, 28 July 2026. Fixture and manifest: content-pipeline/captures/mcp-host-four-scanners/.
Which defect did every tool converge on?
All three findings are the same token at the same line, and the three tools that reached it called it two different things between them.
| Tool | Rule that fired | What it named |
|---|---|---|
| ShipGuard 0.5.1 | hardcoded_secret |
.mcp.json:7, GitHub personal access token, severity high, verdict REVIEW_BEFORE_SHIP |
| Gitleaks 8.30.1 | github-pat |
.mcp.json:7, entropy 4.6841836 |
Semgrep 1.170.0, --config=auto |
none fired | nothing |
| Trivy 0.72.0 | github-pat |
.mcp.json:7, GitHub personal access token |
None of the three tools that fired named the unguarded /mcp_message route or the poisoned tool description, and Semgrep named none of the three seeded defects.
Three tools, two rule names, and the same line of the same file. That convergence is worth more than it looks: three of the four scanners reached the credential with a generic secrets rule, none of them written for MCP. The fourth returned nothing, so this is a check to confirm you have rather than one to assume you have.
What did the grep at the end of that run ask?
Every finding from all four tools was written to one file, one per line, and searched for four strings: mcp_message, the unguarded route path; requireAuth, the middleware it was missing; src/tools.ts, the file holding the poisoned description; and collector.example.invalid, the host that description asks the model to contact.
Four searches, four zeroes. The fixture regenerates and all four scanners re-run every time the script is invoked, so anyone who thinks the result is wrong can produce a different one.
That is not a complaint about the tools, and nothing in the run suggests a bug. A poisoned description matches nothing in a rule corpus. An unguarded route does, and that rule is ours.
Semgrep's zero is narrower than it looks. Our four-tool benchmark put --config=auto over one file holding a hardcoded Stripe live key, an OpenAI key and an AWS access-key pair, and only the Stripe key came back (content-pipeline/benchmarks/RESULTS.md, measured 2026-07-27).
Nothing it applied here fired on a GitHub token in a .json file. We compared secret scanners on their own ground elsewhere, with a fixture built for that question.
What should you check before connecting an MCP server?
Five checks, ordered by how often they catch something rather than by how bad the outcome is. The four-scanner run is already the evidence for row one.
| Check | Why it earns the time |
|---|---|
| No credential inline in the client config | The one MCP defect three of our four scanners reached. An environment reference costs one line. |
| Every tool description read end to end | The description is instructions to the model. It is also the defect no tool in our comparison named. |
| Exact version pinned, diff read on upgrade | postmark-mcp was clean through 1.0.15. Reputation is earned by the releases you already ran. |
| Every route into your MCP handler authenticated | /mcp had the middleware and /mcp_message did not. Both called the same function. |
| The server's own tools listed and justified one at a time | A server installed for one job usually exposes several. The ones you never call are the ones nobody reviews. |
None of the four said anything about rows two, three or five.
How do you stop the fourth row happening again?
Register routes through a single helper that takes the auth middleware as a required argument, so a new route cannot be added without one. It converts a convention into a compile error.
nginx-ui's two registrations were adjacent lines in the same file and still diverged, which is the strongest argument against relying on review for this class of mistake. A reader scanning that file sees two lines that look alike. A type signature sees an argument that is missing.
The advisory's own remediation is the one-line version of this: add the middleware to the route that lacked it. Its second suggestion is the more interesting one, that the empty IP whitelist should deny rather than allow.
Fail-open defaults are what turn a single missing argument into an unauthenticated endpoint. Audit them wherever your own configuration reads "unset" as "everyone", which is exactly the trap Firebase test mode sets in a different corner of the same problem.
Which MCP risks can a scanner reach?
Sorted by what exists as a file, the division comes out clean, and it stays clean across every file scanner we tried.
One MCP connection · two halves
Half of an MCP setup is a file. Half of it is not.
A scanner reads files. The left column is what an MCP connection leaves on disk for it to read. The right column is what the agent meets at runtime, where there is no file to open.
.mcp.json. Right: the tool list, the arguments and the next release, none of which is a file — and none of which any of the four said anything about. Sources: content-pipeline/captures/mcp-host-four-scanners, GHSA-h6c2-x2m2-mwhf, Koi Security on postmark-mcp.The practical consequence is a division of labour rather than a verdict on tooling. Put the left column in a pre-commit hook or a CI step, because a machine re-reads those files on every commit and a person re-reading a config for the ninth time does not. Hand the right column to a person, on the day a server is added and again on the day it is upgraded.
An agent's tool grant is the other half of this picture, and MCP is where that grant now usually comes from.
Nothing in the protocol marks a tool description as untrusted input, and nothing has to. The description is a string a server hands to a client, and the client hands it to a model as part of the context it reasons over. That string is read by the model before it is read by anybody else.
What does the MCP specification actually require?
Revision 2025-11-25 is the current protocol version, and its security document runs to eight named attack sections. The requirements in it are written as normative MUST clauses, which makes them checkable in a way that a list of principles is not.
"MCP servers MUST NOT accept any tokens that were not explicitly issued for the MCP server."
— Model Context Protocol specification 2025-11-25, Security Best Practices
Four of those eight sections did not exist in the previous revision, 2025-06-18: local server compromise, OAuth authorization URL validation, stdio transport in proxy architectures, and scope minimisation. The four that carried over include the confused-deputy problem, where an MCP proxy holding a static client ID toward a third-party authorization server can hand a malicious client a valid authorization code off a consent cookie the user set once, for something else.
What does the spec say about scopes and metadata URLs?
Scope minimisation argues against publishing every scope a server supports and against omnibus scopes such as * or full-access, on the grounds that a stolen broad token then reaches tools the user never intended to expose.
The SSRF section is stranger and more specific. During OAuth metadata discovery an MCP client fetches URLs the server supplies, so a malicious server can point those at 169.254.169.254 and have the client read cloud instance metadata on its behalf.
The spec's advice is to block private and link-local ranges outright, and to apply the same rule to redirect targets. That is a network control rather than a code one, and none of the four scanners we ran here returned anything that touched it.
Why does the client config count as an execution surface?
Starting an MCP server means running a command, and the command lives in the config file.
"If an MCP client supports one-click local MCP server configuration, it MUST implement proper consent mechanisms prior to executing commands."
— Model Context Protocol specification 2025-11-25, Security Best Practices
That requirement is about a command, and the fixture has one — lines 11 and 12 of the generated .mcp.json:
"command": "npx",
"args": ["-y", "postmark-mcp"],
npx -y fetches the package before it runs, so those two lines download and execute code, in the same file the credential was in.
That framing changes what a token in that file means. It is not a setting that happens to hold a secret. It is a credential handed to a process that a configuration file, rather than a human, decided to start, on a machine that also holds your SSH keys.
The same instinct that makes people keep keys out of a Next.js repo applies here. The file is easier to overlook because it looks like editor configuration, and the keys we find most often are the ones somebody filed under settings.
What happens when you fix the one thing a scanner sees?
We made the smallest possible fix. The inline token in .mcp.json became ${GITHUB_PERSONAL_ACCESS_TOKEN}, matching the other server in the same file, which had always supplied its token that way. Nothing else in the tree moved, and the diff below is a real diff -u against the pre-edit copy rather than an illustration of one.

Real capture. @agenticcli/[email protected], 28 July 2026. Fixture and manifest: content-pipeline/captures/mcp-config-fixed-still-poisoned/.
The two registrations the capture prints are these, and they are the fixture's whole reproduction of the advisory:
app.all("/mcp", requireAuth, (req, res) => serveMcp(req, res));
app.all("/mcp_message", (req, res) => serveMcp(req, res));
Why put the unfixed defects in the same frame?
Because a verdict and a repository are different objects, and printing them apart is how the two get confused. SAFE_TO_SHIP is an accurate summary of what the rules matched. It is a poor summary of the tree, and the capture makes that arguable rather than asserted: the unguarded route sits on line 10 of src/server.ts, printed under the verdict, and the injected instruction is quoted from src/tools.ts beneath it.
This is the same shape we found when twelve vibe-coding checks met a real scan, and when an agent holding four destructive tools produced a report identical to one holding a single gated tool.
The two commands printed under the verdict are grep -n mcp_message src/server.ts and sed -n 6,10p src/tools.ts. Both ran against the same tree the scan had just called clean, in the same frame and in the same session, so a reader can re-run either one against the fixture and compare.
Can you trust an MCP server you already reviewed?
In September 2025, Koi Security's risk engine flagged an npm package called postmark-mcp that impersonated the Postmark-maintained repository of the same name. Versions 1.0.0 through 1.0.15 did exactly what they claimed. Version 1.0.16 added a line that BCC'd every outgoing email to an address at giftshop.club.
"These MCP servers run with the same privileges as the AI assistants themselves - full email access, database connections, API permissions - yet they don't appear in any asset inventory, skip vendor risk assessments, and bypass every security control from DLP to email gateways."
— Idan Dardikman, Koi Security, "First Malicious MCP in the Wild", 25 September 2025
Sixteen clean releases is a stronger reputation than most dependencies ever earn, and it was that reputation that made 1.0.16 effective. The developer deleted the package from npm after Koi contacted him, which leaves it running on every machine that had already installed it.
What does pinning a version actually buy?
Not immunity. A pinned version does not stop a compromised release from existing, and it does not help at all if you pin the compromised one.
mcp-remote is the case in point. CVE-2025-6514 scored it 9.6, because connecting that proxy to an untrusted MCP server let the server run operating-system commands on the client machine, in every release from 0.0.5 until 0.1.16 fixed it. What it buys is a moment: the upgrade becomes a decision somebody takes on a particular day, with a diff in front of them, instead of something that happens while a lockfile refreshes on a Tuesday.
That moment is where the review has to land, because it is the only point in the lifecycle where new code and a human attention span are in the same place. Everything before it is trust in a name.
A lockfile is not the control here. A lockfile keeps the build reproducible, which is a different property from keeping it reviewed, and a range that resolves to a new patch release will satisfy every lockfile check you have while installing code nobody read.
Why does the token in our own fixture look the way it does?
The first draft of the fixture filled the 36 characters after ghp_ by repeating the word FAKE. Gitleaks 8.30.1 returned nothing on it, and that is not a coverage gap: the github-pat rule carries an entropy floor, so a placeholder-shaped literal reads as a placeholder.
A comparison published off that run would have measured our own filler. The token now begins ghp_FAKE and fills the rest with high-entropy characters, which is why three of the four tools reach it at all. The correction is recorded in the fixture's header comment and in the capture manifest under fixture_provenance, so the run that was thrown away is on disk beside the run that was published.
The token literal is one line of gen.sh, and compare.sh regenerates the tree and re-runs all four scanners over whatever that line currently says. Anyone who wants to know how much of this result is the fixture rather than the tools can edit one string and find out.
Do MCP clients defend against tool poisoning?
A March 2026 threat model of MCP put seven major clients through the same poisoned-metadata test. Charoes Huang, Xin Huang, Ngoc Phu Tran and Amin Milani Fard applied STRIDE and DREAD across five components of an MCP deployment — host and client, model, server, data stores, authorization server — and then measured what the clients themselves do about a tool description that carries instructions.
"Our analysis reveals significant security issues with most tested clients due to insufficient static validation and parameter visibility."
— Huang, Huang, Tran and Milani Fard, arXiv:2603.22489, 23 March 2026
A client that shows you the tool name and hides the arguments has told you which door opened without telling you what went through it. Our own fixture leans on exactly that gap: the injected text asks the model to read .mcp.json and post it outward, and a client displaying read_project_file alone shows a developer nothing out of the ordinary.
Does a server-side restriction hold on every call?
Arguments are also where a server-side guard can turn out not to be one. Anthropic's Git MCP server takes a --repository flag meant to confine an agent to one checkout, and before 2025.12.18 it never checked repo_path on individual calls against it: CVE-2025-68145.
Where do the paper's defences have to run?
The recommendation it reaches is layered, and none of the layers is a repository scan. The layers are static analysis of the tool metadata, a trace of the model's decision path, behavioural monitoring of what the agent does next, and a client that shows the user more than a tool name.
Every one of those runs where the agent runs. The first already exists outside a client: a scanner that starts your servers and reads their tool lists does it on your machine; judging the result wants a vendor account.
Until a client does it for you, the defence available to a working team is the boring one: read the tool list yourself, before you connect, while it is still short enough to read.
The paper measured clients, not servers. How much of a tool description a person ever lays eyes on is a client decision, and on the seven clients the authors tested it was being made without enough static validation to catch instructions sitting in that description, or enough parameter visibility to show what a call actually carried.
What does ShipGuard not do here?
There is no MCP-specific rule in our corpus at any tier, free or paid. The single result our release gate returned on this fixture came from a generic hardcoded-secret rule that read a .json file without knowing what kind of file it was, and Trivy and Gitleaks reached the same token by the same route. We would rather write that down than let a reader infer coverage from a screenshot.
The unguarded-route rule is the sharper miss, because it exists. It fires on an Express admin route in our own benchmark fixture, and here it stayed quiet on /mcp_message, registered one line below the route that carries the middleware. Two adjacent registrations differing by one argument is a pattern a matcher can hold, and ours does not hold it yet.
What should you ask a vendor selling an MCP scan?
The question that settles it fastest is which of the specification's eight attack sections the product reads, and which of those exist only while the agent is running.
A useful follow-up: ask what the product does on the day a pinned server publishes a new release. That is the postmark-mcp question, and it is answerable by any tool that watches a registry, which is a genuinely different capability from reading your repository.
A third question is cheaper to settle than either. Ask which of our three seeded defects the product would have named, since the fixture, the manifest and the comparison script are all on disk and the run reproduces on any machine with Docker.
The same distinction sits behind auditing what Cursor generates and what Claude Code leaves behind.
Where this leaves a team shipping with MCP today
Run the file half automatically. A secret in a config file, a pinned version in a manifest, an unguarded route in your own source: those are cheap to check on every commit, and three of the four scanners we ran will find the first one for free. The same discipline that keeps Supabase policies honest works here, because both are configuration a machine can read without understanding.
Read the tool list on the day you connect a server, and read the diff on the day you upgrade it. Neither is done by any of the four scanners here, and the one tool that will fetch a tool list for you has to start your servers to do it. Nothing we tested reads an upgrade diff at all, and pretending otherwise is how a clean record through 1.0.15 ends up vouching for 1.0.16.
FAQ
What is MCP security?
What are the most important MCP security best practices?
Can a code scanner detect MCP tool poisoning?
Why did CVE-2026-33032 happen, and what does it teach?
Is a credential in .mcp.json actually a problem?
Does ShipGuard scan MCP servers or tool descriptions?
How do I check an MCP server before I connect it?
What does a green scan on an MCP project actually prove?
────────[ ▮ gate ]────────
Don't ship the next one.
Free, local, no account. Catches this exact bug class before deploy.
$ npx @agenticcli/shipguard scan