AgenticCLI
▮ shipguardis a free, local secure-ship agent for AI-built apps — 5 checks in ~2s.$ npm i -g @agenticcli/shipguard

Slopsquatting: Four Scanners, Two Phantom Dependencies

16 min · jul 2026

Short answer: Slopsquatting is the registration of a package name that a language model invented, so that installing that name fetches somebody else's code. It is not a typo attack, because nobody mistyped anything. We declared two nonexistent dependencies in a real manifest and four scanners returned six findings, not one of them naming either phantom.

On 14 January 2026 someone published a package to npm called react-codeshift. The registry record shows exactly one version, no working code, and a description that reads:

🚫 Placeholder to prevent dependency confusion.

The name was never real. A language model had fused two genuine tools, jscodeshift and react-codemod, into a third that had never existed. By the time Aikido Security's Charlie Eriksen claimed the name himself, it had spread to 237 GitHub repositories, and the reason he claimed it was that downloads were still arriving:

That persistent trickle of 1-4 downloads per day? Those are real. Those are AI agents following skill instructions and triggering npx downloads.

Nothing was breached. A name that described nothing was being requested, every day, by machines acting on instructions no human had read. Somebody friendly got there first.

What is slopsquatting?

The clearest definition sits in the opening line of a 2026 replication study, arXiv:2605.17062:

Spracklen et al. (USENIX Security '25) showed that code-generating large language models hallucinate package names that do not exist on PyPI or npm at rates ranging from 5.2% on commercial models to 21.7% on open-source models, creating an attack surface for slopsquatting -- the registration of malicious packages under hallucinated names.

The word is a blend of AI slop and squatting. The mechanic is older than the word: claim a name people will ask for, then wait to be asked. What is new is who does the asking — an agent working through a task list, at whatever hour, with nobody reading the line it resolves — and how confidently.

Confidence is the load-bearing part. A model that suggests a library does not hedge, and the artifact it hands you looks finished whether or not the pieces inside it exist.

Where does an invented name come from?

react-codeshift is a good specimen because its parentage is visible. Two real tools sit either side of it: jscodeshift, a generic codemod runner, and react-codemod, a set of React-specific transforms. A model that has seen both, asked to modernise a React codebase, produces the average of them. The result is a name that reads like the tool you wanted.

It then escaped through a route nobody was watching. The name entered the world inside a single commit that added dozens of machine-written agent skill files, invoking the package through npx. Those files were forked, copied and even translated, carrying the invented name with them, and none of that copying involved a person reading the line — the review step a pre-ship checklist exists to force.

A hallucinated name in a chat window dies when the window closes. The same name inside a skill file is durable, versioned and reusable, which is exactly what makes an agent framework's install surface worth auditing before you trust it.

Why doesn't a typo defence catch it?

Every defence built against typosquatting starts from one assumption: that a human meant one thing and their fingers produced another. Somebody means to install requests, types reqeusts, and collects whatever an attacker registered under the misspelling. Edit-distance checks, warnings on similar package names and registry-side blocks on near-misses all follow from that premise.

None of that machinery has anything to grip here. The developer did not slip, because the developer did not type. A model produced a name, an editor accepted it, and it landed in the manifest looking as deliberate as the line above it.

react-codeshift is not a misspelling of anything. It is a competent-looking synthesis of two real names, which is precisely why nobody looked twice, and a filter tuned to catch near-misses of popular packages sees nothing unusual in it at all.

How far from a real name is an invented one?

Measured against a comparison set whose every candidate was confirmed at the registry before anything was compared to it, react-codeshift sits five edits from react-codemod and six from jscodeshift: a sibling of both, a mistyping of neither.

The two names we invented for our own fixture behave the same way. stripe-webhook-verify is fifteen edits from stripe; flask-request-validator-lite is five from flask-request-validator, and five is the closest any invented name in this fixture gets to one a registry confirmed real. The misspelling this section opened with, reqeusts for requests, is two. Script and measurements: content-pipeline/captures/slopsquat-registry-check/edit-distance.py, run 2026-07-28.

A misspelling carries evidence of the error inside itself. An invented name is only wrong relative to a fact that lives somewhere else entirely — in a registry, on the day you ask.

What do four scanners find on a manifest with two phantom dependencies?

We built the smallest honest test of that: a service of the shape an agent writes in one sitting, with an Express webhook receiver, a Flask reporting endpoint, and five declared dependencies across package.json and requirements.txt. Two of the five have no package behind the name.

The four tools are the ones already sitting in most pipelines, and they were picked for coverage of the three shapes a manifest defect can take: a secrets matcher, a pattern matcher over source, and a dependency scanner that parses the manifest itself. If any category was going to speak, it was the last one.

The tree also carries a hardcoded Stripe test key, deliberately. Without something the scanners are known to catch, a clean report reads as a broken harness rather than a result, and we have watched these same scanners behave very differently on secrets. The run has to show them working before it shows them silent.

What came back?

Real ShipGuard 0.5.1 terminal output from capture slopsquat-four-scanners — 5 files read and 6 findings from 4 scanners run over a manifest listing 5 declared dependencies, 2 of them with no package behind the name; the per-tool split is ShipGuard 2, Gitleaks 1, Semgrep 1, Trivy 2, and none of the six names a nonexistent dependency

Real capture. @agenticcli/[email protected], Gitleaks 8.30.1, Semgrep 1.170.0, Trivy 0.72.0, 28 July 2026. Fixture and manifest: content-pipeline/captures/slopsquat-four-scanners/.

Four of the six findings are the same key at line 8, caught independently by three of the tools. Semgrep contributed an Express CSRF-middleware audit rule and no secrets rule, which is a known property of its open-source distribution rather than a miss.

ShipGuard's own verdict on the tree was REVIEW_BEFORE_SHIP, earned entirely by the credential. Strip that one line out and the same tree scans clean, with both invented dependencies still sitting in the manifest.

Tool Findings Naming a nonexistent dependency
ShipGuard 0.5.1 2 0
Gitleaks 8.30.1 1 0
Semgrep 1.170.0 1 0
Trivy 0.72.0 2 0

Why is Trivy's row the sharp one?

Of the four, Trivy is the one whose entire job is dependencies, and on this tree it did that job: it opened requirements.txt, resolved the pinned Flask release to a published CVE, and reported it. On the very next line of the same file sat a project that does not exist, and Trivy said nothing.

That is not a coverage gap, and no amount of rule-writing closes it. An advisory database can only carry records about packages that were published, reviewed and found wanting. A name nobody ever registered has no advisory, no maintainer, no release history and no CVE, so it is indistinguishable from a package that is simply fine.

The absence of a warning is the absence of a record. It is the same failure mode as a scanner reading a workspace it cannot execute: the tool is answering accurately, and the question it answers is not the one you had.

What happens when nobody has registered the name?

One invented name · two owners

The same package.json, resolved on two different days

Nothing about the code changes between these two paths. The model wrote the same name, the commit landed the same way, the install command is the same command. What differs is whether anyone had claimed the name by the time the resolver asked.

PATH A · NOBODY CLAIMED IT 1 · THE MODEL WRITES A NAME stripe-webhook-verify Plausible. Nothing behind it. 2 · IT REACHES A MANIFEST package.json README skill file Copied forward by whoever reads it next. Nobody looks it up. 3 · SOMEBODY RUNS INSTALL npm install The resolver asks the registry. 4 · THE REGISTRY ANSWERS 404. No package under this name. The install stops here. 5 · RESULT Nothing is installed. You find out in the same second. The name behaved like a typo. PATH B · SOMEBODY CLAIMED IT 1 · THE MODEL WRITES A NAME same as A stripe-webhook-verify The model is not what changed. 2 · IT REACHES A MANIFEST package.json README skill file Same file. Same commit. Still nobody looks it up. 3 · SOMEBODY RUNS INSTALL npm install Same command. Same question. 4 · THE REGISTRY ANSWERS 200. Somebody registered the name. Install scripts run. npm Docs. 5 · RESULT The install succeeds. No file in the repo changed. Nothing new for a scanner to read.
One invented package name down two paths. Left: no package exists under it, the resolver returns 404, and the install fails in the loudest possible way — the lucky outcome, which is why this panel is the green one. Right: somebody registered the name first, the resolver returns 200, and npm's own documented lifecycle runs the package's install scripts as part of the install. The model, the commit and the command are identical on both sides. The only variable is who owned the name on the day the resolver asked.

The figure holds one name and two outcomes. Our own manifest produced both at once: read on 28 July 2026, it sent stripe-webhook-verify down the left path and react-codeshift down the right. Nothing in the two lines of the file tells them apart.

The left path is the forgiving one. A resolver that cannot find a name stops, prints an error and installs nothing, so the invented dependency behaves like a typo: annoying, visible, fixed in a minute.

Once the name is claimed, the same manifest resolves cleanly. npm's own documentation places the install lifecycle scripts "after the actual installation of modules into node_modules, in order, with no internal actions happening in between", so code the publisher wrote runs during the install itself.

Microsoft's Defender research team observed the same mechanic in a live campaign, writing that "The malicious code executes the moment a victim runs npm install; no require() from victim code is needed." Installation is execution, and it happens before your own code has run a single line.

What does --ignore-scripts actually buy you?

npm's install reference states that with the flag set, "npm does not run scripts specified in package.json files", which removes the automatic half of the problem and closes the path from resolution to execution without any decision of yours.

What it does not do is make the dependency safe. The package is still on disk, still in your tree, and still executes the moment anything imports it. Some legitimate packages genuinely need their install step, so a blanket flag turns into a list of exceptions you maintain, and every exception is a place the reasoning has to be redone.

Run it anyway, and treat the list of packages that broke as its own finding. A dependency that insists on running code during installation has told you something about itself, which is more than most manifests give you.

What do the registries say about these five names?

The check the four scanners never make takes one request per name. We parsed the names out of the same fixture's manifests rather than typing them, and asked each registry directly: registry.npmjs.org for the npm names, and for the Python ones the JSON document PyPI publishes for a project.

Real terminal output from capture slopsquat-registry-check — 5 dependency names checked against their registry, 2 return 404, and 1 resolves to a placeholder published to block a squat, with each npm answer showing its creation date, version count and description

Real capture. curl against registry.npmjs.org and pypi.org, 28 July 2026. Probe and manifest: content-pipeline/captures/slopsquat-registry-check/.

Nothing sits behind either invented name: no maintainer, no version history, no download count, because there was never a package there to have one. A resolver asking for either gets the same dead end a misspelling gets.

Is "does the package exist" enough?

The react-codeshift row comes back 200, so a gate that stops at the status code waves it through having learned nothing at all about it.

What the record actually says is that the name was created this January, carries a single published version, and describes itself as a placeholder against dependency confusion. That is a defender's claim on a name a model invented, and it happens to be safe. The next name in that position may not be.

So read the fields. A dependency your editor suggested that was created last month, has one version, and cannot show you a repository or a download history is not automatically hostile. It is unestablished, which is a different thing and a fair reason to stop and look, in the same way an unfamiliar route appearing in a diff is a reason to read it rather than a reason to panic.

How often does a model invent a name?

Two studies have measured this, and neither of them is ours. Spracklen and colleagues, published at the 2025 USENIX Security Symposium and available as arXiv:2406.10279, generated 576,000 code samples across 16 models and found hallucinated packages in at least 5.2% of commercial-model output and 21.7% of open-source-model output, across 205,474 distinct invented names.

The 2026 replication narrowed the band considerably. Across 199,845 prompts and five newer models it measured 4.62% to 6.10%, a compression rather than a disappearance. It is a single-author preprint that has not been peer-reviewed, and it should be read as early evidence rather than settled fact.

Both studies validate their output against the live PyPI and npm name lists, which is why a hallucination is a checkable event here rather than a judgement call. A name either resolves or it does not, and that is the same binary a reader can run themselves.

Can you outrun this by switching models?

The same preprint answers that, and the answer is the least comfortable line in it:

Beyond replication, we identify a set of 127 package names (109 on PyPI, 18 on npm) that all five evaluated models invent identically

Following coordinated disclosure, 53 of those names (41 on PyPI, 12 on npm) were still registrable by an attacker after each registry's existing defenses.

Four vendors and five training runs produced the same inventions. A defence built on choosing a better model assumes the models fail independently, and on this measurement they do not. It is the same reasoning that makes the harness around a model matter more than the model.

Told exactly which names were at risk, the registries still left 53 of the 127 claimable, because a registry cannot refuse a name for being plausible. Refusing every name a model might invent means refusing most reasonable names.

Does this stop at package names?

A July 2026 preprint, arXiv:2607.07433, extends the same trick from packages to anything an agent resolves by name: a repository to clone, a skill to install. The authors call it adversarial hallucination squatting, and their method is to compute which names a model is most likely to invent around a trending resource, then register those first.

We empirically demonstrate that hallucinated resource generation occurs at high rates, up to 85% in repository cloning scenarios and up to 100% in skill installation, and that these hallucinations transfer between foundational models and different prompts.

Read alongside the react-codeshift incident, where the invented name travelled inside skill files rather than inside code, that is the shape the problem is taking. The dependency manifest was only the first place a machine-written name got resolved with no human in between, and the same question now applies to every MCP server an agent connects to.

What can you check before you ship?

The five below are ordered by what they cost you rather than by how much they catch, and none of them needs a vendor.

Check How
Does every declared name exist? One registry request per name, in CI
Is a newly added name established? Read creation date, version count, repository
Did install scripts run? npm install --ignore-scripts, then audit what needed them
Is the tree pinned? A committed lockfile, reviewed in the diff
Did a human read the diff's new dependencies? Treat an added manifest line as a code change

The first is cheap, deterministic and absent from every tool we ran. The last is the one that actually holds: a new dependency line is a decision to run somebody else's code on your machine, and it deserves the attention a pasted credential gets in review.

Ordering matters more than completeness. Existence is a yes-or-no question a machine settles, so it belongs in CI; establishment is a judgement, so it belongs to whoever approves the diff.

What this doesn't cover

Two limits, stated plainly, because an account that lists only what it proved is worth less than one that marks its own edges.

We did not measure a hallucination rate ourselves. Doing that honestly means the scale the two cited studies worked at, and a run of a few dozen prompts dressed up as first-party evidence would be worse than a citation. The rates here belong to their authors, with their methods and their caveats, and one of those two papers has not been through peer review.

We also did not register a name to demonstrate the attack. Claiming a plausible invented name, even defensively, means publishing to a public registry, and the incident that opens this piece is somebody else's real registration that anybody can go and read.

And what our own tool does not do

ShipGuard does not answer this question. It read five files on our own fixture, found the key we planted, and said nothing about either phantom dependency, which is right for a gate that reads the code you are about to ship.

Asking whether a name exists means asking a registry over the network: a different question, with a different failure mode, and no rule in our corpus at any tier addresses it.

# npm — 200 means a package exists under this name, 404 means nothing does
curl -s -o /dev/null -w '%{http_code}\n' "https://registry.npmjs.org/$NAME"

# PyPI — the same question, asked of the project's JSON document
curl -s -o /dev/null -w '%{http_code}\n' "https://pypi.org/pypi/$NAME/json"

Both are single unauthenticated GETs, and neither needs an account. The probe that parses the names out of package.json and requirements.txt first, so the check cannot drift from the manifest it claims to be checking, is content-pipeline/captures/slopsquat-registry-check/probe.sh.

The shell above does answer it, it costs nothing, and that is a better thing to hand you than a roadmap. Publishing the check we do not sell is cheaper for us than the alternative, which is letting a reader assume a scan covered something it never looked at.

FAQ

What is slopsquatting?
Slopsquatting is the practice of registering a package name that a language model invents, so that anyone who installs that name gets the registrant's code instead of an error. The preprint arXiv:2605.17062 defines it in one line as the registration of malicious packages under hallucinated names. The name blends AI slop with squatting. It differs from typosquatting in who makes the mistake: typosquatting waits for a human to mistype a real name, while slopsquatting waits for a model to invent a name nobody typed at all, then claims that name before anyone checks whether it was ever real.
How is slopsquatting different from typosquatting?
Typosquatting depends on a human slip. Somebody types reqeusts instead of requests, and an attacker who registered the misspelling collects the install. Slopsquatting needs no slip, because the developer never typed the name. A model wrote it, an autocomplete accepted it, and it landed in a manifest looking exactly as deliberate as the line above it. That difference matters for defence: spell-checking a dependency name against a list of popular packages catches the first attack and does nothing about the second, because a hallucinated name is usually not a near-miss of anything.
Do security scanners detect hallucinated package names?
Not in the run published here. AgenticCLI built a service declaring five dependencies, two of which have no package behind the name, and read it on 28 July 2026 with @agenticcli/shipguard 0.5.1, Gitleaks 8.30.1, Semgrep 1.170.0 and Trivy 0.72.0. The four tools returned six findings between them and none named a nonexistent dependency. Four of the six were the same hardcoded test key, so the tools were working. Trivy parsed requirements.txt, resolved the real pin to a published CVE, and returned nothing for the phantom on the following line.
How do I check whether a package an AI suggested is real?
Ask the registry before you install. For npm, request the package document at registry.npmjs.org and read the status code: 200 means a package exists under that name, 404 means nothing does. For Python, the equivalent is the PyPI JSON endpoint for the project. Both are single unauthenticated requests and neither needs an account. Then read what came back rather than stopping at the status: the creation date, the number of published versions and the description tell you whether you found an established library or a name somebody claimed last week.
Does an install fail safely if the package does not exist?
Yes, and that is the outcome you want. A resolver that gets a 404 stops, prints an error and installs nothing, so the invented name behaves exactly like a typo and you find out immediately. The dangerous case is the opposite one. Once somebody has registered the name, the resolver succeeds, and per npm's own documentation the preinstall, install and postinstall scripts run as part of the install. Microsoft's Defender research team describes the same mechanism in a live campaign: the code runs at install time, with no import from your own code needed.
How often does an AI model invent a package name that does not exist?
The measured figures come from two studies rather than from us. Spracklen and colleagues, published at the 2025 USENIX Security Symposium, generated 576,000 code samples across 16 models and found hallucinated packages in at least 5.2% of commercial-model output and 21.7% of open-source-model output, spanning 205,474 distinct invented names. A 2026 single-author preprint, not yet peer-reviewed, replicated the method on five newer models over 199,845 prompts and measured a much narrower band of 4.62% to 6.10%. The spread compressed. The floor did not reach zero.
Can I just switch to a better model to avoid this?
The evidence says no, and this is the least obvious result in the literature. A 2026 replication preprint tested five frontier models from four different vendors and found 127 package names that every one of them invents identically, of which 53 were still registrable by an attacker after coordinated disclosure to PyPI Security and Socket.dev. Different companies and different training runs produced the same inventions. A defence built on picking a better model assumes the models fail independently, and on this measurement they do not.
Does ShipGuard flag a dependency that does not exist?
No. On the fixture in this article ShipGuard read five files and every finding it reported was the hardcoded test key, and it said nothing about either phantom dependency. That is the correct result for what it is: a release gate that reads the files you are about to ship. Deciding whether a name exists means asking a package registry over the network, which is a different question with a different failure mode, and there is no rule for it in the corpus at any tier. The check that answers it is one unauthenticated request per name against the registry, and it needs no account and no vendor.

────────[ ▮ gate ]────────

Don't ship the next one.

Free, local, no account. Catches this exact bug class before deploy.

$ npx @agenticcli/shipguard scan