AI Blog

Skill Scanners Read a Different File Than the Agent Runs

Trail of Bits bypassed the detectors on three skill-distribution platforms in June, a July study packed 1,613 malicious skills past all eight scanners tested, and on 17 August OWASP gave poor scanning its own entry in the first Agentic Skills Top 10. The scanner inspects a file at rest; the agent constructs a program from it at run time — and the attacker picks where the two disagree.

By Agentic AI Wiki 15 min read

A skill scanner and the agent that installs the skill are not looking at the same artefact, and the attacker gets to choose where they differ. Trail of Bits walked past the detectors on three skill-distribution platforms in June — three of its four bypasses took under an hour each to build — and a July study packed 1,613 real malicious skills so that every one of eight scanners missed them more than ninety per cent of the time. On 17 August OWASP wrote the conclusion into a standard, giving poor scanning its own entry in the first Agentic Skills Top 10. The green checkmark on a listing is not a weak signal about your skill; it is a confident signal about a different file.

What actually landed

Three results, published across eleven weeks, that say the same thing from three directions.

DateResultScope
3 June 2026 Trail of Bits bypasses marketplace skill detectors ClawHub (VirusTotal plus a guard model), Cisco AI Defense's scanner, and the three scanners wired into Vercel's skills.sh. Four distinct bypasses; three built in under an hour each.
10 June 2026 Cloud Security Alliance research note, AI Agent Skill Scanners: Bypassed Across the Board Eight open-source scanners tested against real attack samples.
July 2026 Cloak and Detonate (SkillCloak), arXiv 2607.02357 1,613 in-the-wild malicious skills, eight scanners, two production agents. Self-extracting packing bypasses every scanner over 90% of the time; the payloads still execute correctly.
17 August 2026 OWASP Agentic Skills Top 10 v1.0 Ten risks in the skill layer across OpenClaw, Claude Code, Cursor, Codex and VS Code — including AST07 update drift and AST08 poor scanning.

Two of those are adversarial research, which always reads worse than the average case. The OWASP entry is the part that matters commercially: a standards body has now written down, as a named risk, that the control most marketplaces sell as their safety story does not hold. Everything downstream of that — procurement questionnaires, vendor security reviews, the "is it scanned?" box on an internal approval form — is now asking a question with a known answer.

The scanner reads a file; the agent reads a program

A skill bundle read two different ways One published skill bundle on the left contains a SKILL.md file, a Word document in an assets folder, and a compiled Python bytecode file. The scanner path along the top reads only the first slice of the plain text, cannot open the archive or the bytecode, and returns a clean verdict. The agent path along the bottom renders the whole instruction file, unpacks the archive, imports the bytecode, and fetches a further payload at run time, so the malicious code executes. The two readers disagree about what the same bytes contain, and the attacker chooses where they disagree. The published bundle SKILL.md assets/brief.docx lib/format.pyc one artefact, two readers What the scanner reads the first slice of plain text — truncated by a token budget no archive, no bytecode one verdict, produced once What the agent reads the whole instruction file unzips the document it is told to imports the compiled module fetches whatever step three names every session, after every update Verdict: clean a badge on the listing page Payload runs credentials, source, backdoor semantics fully preserved THE GAP THE ATTACKER PICKS
Same bytes, two readings. The attacker only has to find one place where they diverge.

Every published bypass is a variation on one structural fact. A skill is not a binary with an entry point; it is a bundle of instructions that only becomes a program when a model reads it and decides to act. The scanner examines the bundle at rest. The agent constructs something else from it at run time — following its instructions to open a file, import a module, fetch a URL, or read a document — and it is that constructed thing that has behaviour. The gap between the two is a parser differential, and parser differentials have been the reliable way to defeat content inspection since the first email virus scanner.

The reported techniques fall into three groups, and none of them is novel research:

  • Make the scanner stop reading. Trail of Bits prepended roughly a hundred thousand newlines to the malicious section of a skill file. The scanning harness truncated the oversized content before handing it to its guard model; the payload sat past the cut. The agent, reading the same file for its own purposes, had no such limit.
  • Put the payload somewhere the scanner does not parse. An innocuous instruction file telling the agent to extract its real instructions from a .docx — a ZIP of XML with a shell script inside. Or shipping legitimate Python source next to poisoned compiled bytecode. In both cases the text the scanner reads is clean, because the text the scanner reads is not the payload.
  • Do not include the payload at all. This is the SkillCloak result and the strongest of the three: self-extracting packing keeps the malicious component out of the install-time view entirely and reconstitutes it during execution. There is nothing for a static reader to find, which is why it clears every scanner tested at over ninety per cent.
Scanner bypass rate by evasion technique Horizontal bar chart of measured bypass rates against agent skill scanners. Self-extracting skill packing exceeds ninety per cent against every one of the eight scanners tested. Structural obfuscation reaches ninety-six per cent against the single hybrid scanner and exceeds eighty per cent against most purely static scanners. A separate hand-built study, noted below the chart, bypassed the detectors on all three marketplace platforms it tested, with four distinct bypasses. Bypass rate against agent skill scanners 25% 50% 75% 100% Self-extracting packing every one of 8 scanners >90% Structural obfuscation vs. the hybrid scanner 96% Structural obfuscation vs. most static scanners >80% SEPARATELY, BY HAND: EVERY MARKETPLACE DETECTOR TESTED WAS BYPASSED four bypasses across three platforms; three of the four took under an hour each to build
Measured against 1,613 real malicious skills. The packing result is the one that does not have a patch.

The detail worth dwelling on from the SkillCloak paper is not the bypass rate. It is that the cloaked skills, run through Claude Code and Codex, still did exactly what the uncloaked versions did. Evasion cost the attacker nothing in capability — which means there is no rate-limiting trade-off to lean on, no "the stealthy version is a weaker attack." The equilibrium is that everything malicious is packed.

Putting a model in the scanner makes the scanner a target

The industry's answer to "static analysis cannot read prose" was to put a model in the scanner. That fixes the expressiveness problem and introduces a worse one: the scanner is now a component whose input is written by the adversary and whose output is a security decision. That is the definition of the exposure in prompt injection, only pointed at your control plane rather than your agent.

It works about as well as you would expect. In the Trail of Bits work, one bypass simply framed a malicious package-registry configuration as a routine enterprise requirement; the LLM scanner accepted the framing and downgraded the severity to low. No obfuscation, no encoding — the scanner was argued out of its finding, in the same plain English the skill was written in.

This produces a genuinely unpleasant property. A static scanner fails closed-ish when it does not understand something; a model-based scanner produces a confident, well-written verdict either way, and the verdict is graded on the same channel it is being attacked through. Adding a second model does not fix it, and the SkillCloak numbers bear that out — the hybrid scanner was the one structural obfuscation beat 96% of the time, not the least effective of the group. If you are wondering what a poisoned description does to a tool-selection decision, the same shape is dissected in MCP tool poisoning.

Even a perfect verdict expires before you use it

Grant, for a moment, a scanner that reads exactly what the agent reads and cannot be talked to. It still only tells you about the version it read. OWASP's AST07 names this directly as update drift: a skill that changes after approval invalidates the review that approved it, and most installs track a name rather than a revision.

This is ordinary time-of-check-to-time-of-use, and package ecosystems solved it a decade ago with lockfiles and content hashes. Skills largely have not, because the distribution layer grew out of "paste this into your agents folder" rather than out of a package manager — the point the Agent Plugins 1.0 format left unaddressed when five vendors agreed on a directory layout and explicitly declined to agree on install, provenance or permissions. The registry landscape has the same hole, examined index by index in the four MCP indexes.

The practical version: if you cannot state which bytes you approved, you did not approve anything. That is one hash per skill and one file in the repository, and it is the cheapest item on this page by a wide margin.

The checkmark is the actual product defect

A scanner that misses things is a tool with a false-negative rate, and every security tool has one. That is not the complaint. The complaint is what the marketplace does with the result: it converts unknown into approved, publishes the conversion next to a download button, and thereby moves the risk onto the person least able to evaluate it.

The base rates make this concrete. When Koi Security audited ClawHub on 1 February 2026 it reported 341 malicious entries out of 2,857 skills — roughly one in eight — during a campaign that eventually accounted for over a thousand malicious listings. In a population like that, a badge whose false-negative rate is above ninety per cent for packed payloads is not a weak filter. It is close to a random relabelling of a heavily poisoned registry, presented as due diligence.

Three places to put the control on a third-party skill Three columns comparing where a control on an installed skill can sit. An install-time scan is cheap and produces an inventory, but it reads a different artefact than the agent runs and expires the moment the skill updates. A content pin plus a review of what the skill may reach fixes the time-of-check problem and costs one hash and one manifest. Run-time capability and egress limits are the only control that binds what the payload can do rather than what the file appears to say, and cost a real boundary in the runtime. Scan at install reads the file as published verdict expires on first update can itself be talked to WORTH KEEPING AS an inventory, not a gate Pin and re-review install by content hash updates re-enter review closes check-versus-use drift COSTS one hash, one lockfile Constrain at run time binds what the payload can do no filesystem, no egress by default indifferent to how it was packed COSTS a real boundary in the runtime
Only the right-hand column constrains behaviour rather than appearance.

Which points at the only control in the row that survives all three failures at once. A capability boundary does not care whether the payload was packed, whether it arrived in a .docx, or whether the file changed since Tuesday, because it binds what the code can reach rather than what the file appears to say. No filesystem outside the working directory, no credential the model can name (see scoped credentials), and no network egress except to an allowlist — the discipline in egress control for agents. A skill that cannot reach your cloud metadata endpoint is not made safe by a scan and is not made dangerous by the absence of one.

The steelman

Deleting the scanner would be the wrong lesson, and two arguments for keeping it are good.

First, the base rate cuts both ways. The ClawHub campaign was not subtle — long documentation files with a "prerequisites" section telling the user to paste a command into a terminal. Scanners catch that, and catching the unsophisticated majority of a poisoned registry is real value even at a terrible rate against a determined attacker. Second, a scanner produces the artefact that the governance layer actually needs: a list of what is installed, what it declares, and what changed. That is agent inventory, and almost nobody has it.

Keep the scan. Change what it outputs. A scan that emits a row in an inventory and a diff against last week's version is doing honest work; the same scan emitting a green badge on a listing page is making a claim it cannot support.

What to actually do

If your team installs third-party skills

Pin by content hash and make an update a review event rather than a silent fetch. Then ask the question a scan cannot answer: if this skill were malicious, what could it reach? Answer it for filesystem, credentials and network separately, and fix whichever answer is worst. That is a thirty-minute exercise per skill and it dominates any amount of scanning, for the reasons laid out in agent supply-chain security.

If you run an agent platform

Put the boundary under the skill, not in front of it. Default-deny egress and a working-directory-scoped filesystem turn skill vetting from a detection problem into a blast-radius problem, and blast radius is something you can actually bound — see sandboxing and safe execution. Treat the scanner's output as telemetry feeding that boundary's policy, never as an admission decision on its own.

If you operate a marketplace

The badge is the liability. Publish what was checked, which revision it was checked against, and when — a dated, versioned statement of scope that a reader can reason about. Publish provenance and a hash that installs can pin to. Anything that renders as a binary safe/unsafe judgement on a mutable artefact will keep being wrong in the direction that hurts your users.

The durable principle: a control that inspects an artefact is only as good as the agreement between your parser and the real interpreter's, and for agent skills there is no agreement at all — the real interpreter is a model that will follow instructions to fetch, unpack and execute things your scanner never saw. Scan for inventory and for the unsophisticated majority. Pin so that what you approved is what runs. But put the control that you actually rely on at run time, where it binds capability instead of appearance, because that is the only layer where the attacker does not get to choose what you are reading.

FAQ

Should we stop using agent skill scanners?

No — stop treating their output as an admission decision. They reliably catch unsophisticated malicious skills, which are the majority in a poisoned registry, and they produce the inventory and change-diff that your governance layer needs. What they cannot do is certify a skill as safe, and any workflow that reads a clean verdict as approval is relying on the one thing the research says does not hold.

Why does adding an LLM to the scanner not fix this?

Because it changes the scanner from a parser into a participant. The model reads text written by the adversary and emits a security verdict, so the scanner inherits the whole prompt-injection exposure — in the published work, one bypass simply framed a malicious configuration as a corporate requirement and the scanner downgraded it to low severity. The measured evidence points the same way: structural obfuscation performed best, at 96%, against the hybrid scanner rather than against the purely static ones.

What is self-extracting packing, exactly?

The malicious component is not present in the form that gets installed and reviewed. What ships is a benign-looking bundle that reconstructs the payload during agent execution. A static reader has nothing to find because at the moment it looks, there is nothing there — which is why it clears every scanner tested more than ninety per cent of the time, while the reconstituted payload behaves identically to the original.

Does this apply to MCP servers as well as skills?

The same structure, with the roles moved. An MCP server ships code you run and tool descriptions your model reads, so both the install-time artefact and the run-time text are attacker-controlled surfaces, and a description can be rewritten by the server after you approved it. Pinning and run-time confinement are the same answer; the tool-description half is covered separately under MCP tool poisoning.

Is a content hash enough on its own?

It is necessary and not sufficient. A hash guarantees that what runs is what you reviewed; it says nothing about whether your review was any good, and the research here is precisely that the review is unreliable. Pin so the verdict stops expiring, then spend the effort on capability limits, which are the part that holds even when the review was wrong.

Further reading

On this wiki:

Sources: