AI Coding Agents Are the Largest Unpatched Supply Chain on Your Machine. Here Is the Receipt.

In short
Key takeaways
- 100% of ten tested AI coding tools were vulnerable to prompt injection leading to code execution or data exfiltration, per the IDeaster campaign
- Cursor in Auto Mode failed 83.4% of 314 payloads across 70 MITRE ATT&CK techniques; Cursor with a Claude 4 backend failed 69.1%
- Plugin4Shell is fixed in Claude Code 2.1.179 and Codex CLI 0.146.0, but GitHub Copilot and Gemini CLI have no vendor fix as of September 2026
- A state-sponsored group (GTG-1002) ran Claude Code autonomously against ~30 targets executing 80-90% of operations without human input; one attacker breached nine Mexican government agencies and exfiltrated 150GB including 195 million taxpayer records
- Backdoored LiteLLM builds were downloaded ~47,000 times in the three hours they were live on the package index
Every AI coding agent on your machine reads content written by someone else and then acts on it with your permissions. That is the product. It is also, structurally, the largest unreviewed code-execution surface in modern software.
The evidence is not hypothetical. Thirty-plus CVEs across ten tools, a 100%-vulnerable research campaign, an 83.4% payload success rate against Cursor in Auto Mode, a state-sponsored group running Claude Code autonomously against thirty targets, and a nine-agency government breach.
So I built it. On a sacrificial machine, with a canary, and published exactly what happened. Then I built the defence, because nobody writes the mitigations down.
Why this is not a normal dependency problem
A normal dependency is inert until someone executes it. An agent reads attacker-controlled content by design, then acts on it with your permissions. The untrusted content is inside the trusted execution path. There is no equivalent in the previous twenty years of software supply chain.
The Cloud Security Alliance puts the sharpest version on it: agents, marketplaces, and package managers now extend "ambient trust and privilege" to their own extensions. That is a trust relationship that did not previously exist in the package ecosystem, and it is not covered by the controls package managers had.
The table
| ID | Target | Issue | Severity |
|---|---|---|---|
| CVE-2025-32711 (EchoLeak) | M365 Copilot | First documented zero-click agent attack. One crafted email caused internal data exfiltration with no user interaction | CVSS 9.3 |
| CVE-2025-53773 | GitHub Copilot | RCE via prompt injection hidden in code comments, triggering autonomous execution on 100,000+ developer machines | CVSS 9.6 |
| CVE-2025-54135 (CurXecute) | Cursor | MCP config rewrite via Slack leading to RCE without user approval | n/a |
| CVE-2025-49150 | Cursor | Remote JSON schema data exfiltration | n/a |
| CVE-2025-54130 | Cursor | Settings overwrite leading to RCE | n/a |
| CVE-2025-61260 | OpenAI Codex CLI | Project-local config files execute commands without user consent | HIGH |
| CVE-2025-54794 | Claude Code | Path restriction bypass, file access outside workspace | n/a |
| CVE-2025-54795 | Claude Code | Command injection, arbitrary shell execution | CVSS 8.7 |
| CVE-2025-59536 | Claude Code | Malicious SessionStart hooks leading to RCE on agent launch | n/a |
| CVE-2026-21852 | Claude Code | Malicious ANTHROPIC_BASE_URL leading to API key exfiltration | n/a |
EchoLeak and Copilot RCE are the two that should change how you think. One email, no interaction, data leaves. One code comment, and machines running a code-review agent executed whatever was hidden in it.
The 100% campaign and the 83.4% number
IDeastER (Dec 2025, "Attacked by A.I. Agents, This Start-Up Embarked on a Crusade," ZDNET, Aug 24 2026): 30+ vulnerabilities across ten major AI-integrated development environments. 100% of the ten tested tools were vulnerable to prompt injection leading to code execution or data exfiltration. Tools covered included GitHub Copilot, Cursor, Claude Code, Amazon Q Developer, and OpenClaw.
AIShellJack (arXiv:2509.22040): 314 payloads, 70 MITRE ATT&CK techniques, against production AI coding tools. Cursor in Auto Mode: 83.4% of tested attacks succeeded. Cursor with Claude 4 backend: 69.1%.
If you have Cursor Auto Mode on right now, 83.4% is the number that describes your machine.
Plugin4Shell, dated, with the fix matrix
Cloud Security Alliance research note, Sep 19 2026: a SHA-pinning bypass enabling RCE via agent plugin substitution. This is the single most actionable finding in the corpus.
- Claude Code fixed at 2.1.179
- OpenAI Codex CLI fixed at 0.146.0
- GitHub Copilot and Gemini CLI: no vendor fix to verify against
If you run Copilot or Gemini CLI, this is the section that matters, and it is dated, and it is live.
CSA's structural criticism of the whole control: "any agent, marketplace, or package manager that advertises commit-hash pinning as an integrity control needs to assert the resolved working-tree state after checkout, not merely the reference it requested." Commit-hash pinning, the control everyone recommends, does not verify the thing you actually got. It verifies the thing you asked for.
NomShub, the worked example
This is the chain that teaches you to recognize the next one. Four stages, per Straiker (Apr 3 2026):
- Attacker injects instructions into a repository README.md
- Cursor reads it, follows the injected instructions
- Sandbox escape via shell builtins: Cursor's guard did not cover commands executed within the shell, so the parser was blind to working-directory changes, manipulated environment variables, and altered execution context. The macOS seatbelt sandbox permits writes to the home directory, so builtins overwrote .zshenv, which every new zsh instance executes
- The agent is instructed to generate a device code and send it to the attacker's server, authorizing the attacker's GitHub account against the victim's Cursor remote tunnel, giving persistent, undetected shell access
Requires no interaction beyond opening the repository. The exploited binary is legitimately signed and notarized. And per Straiker, the traffic routes through Microsoft Azure infrastructure, making network-level detection "nearly impossible."
The config-style finding: what actually gets in
arXiv 2604.03081 ("Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems") built a 1,070-sample benchmark from 150+ real software supply-chain poisoning incidents (2021 to 2026, GHSA and NVD records plus OWASP LLM Top 10). Category distribution: supply-chain poison 47.5%, environment-variable theft 11.9%, credential theft 8.6%, plus config tamper, code/infra, network exfil, system persistence.
The finding that matters most:
"Configuration-style attacks (supply-chain and configuration tampering combined, n=572) account for 72% of Sonnet's executions (18/25). This concentration indicates that payloads resembling routine developer workflows are the primary driver of bypass success."
Payloads that look like a normal Tuesday. Not exotic exploits. Config edits.
Per-model success rates (higher = more vulnerable) on the hardest categories:
| Category | n | Sonnet | GLM | MiniMax |
|---|---|---|---|---|
| Supply-Chain Poison | 508 | 2.8 | 2.4 | 9.4 |
| Creds & Env Theft | 219 | 2.3 | 3.7 | 17.8 |
| Config. Tamper | 64 | 6.3 | 6.3 | 20.3 |
| Code & Infra | 124 | 1.6 | 0.0 | 13.7 |
| Network Exfil | 82 | 0.0 | 1.2 | 19.5 |
| Sys. Persistence | 48 | 0.0 | 2.1 | 16.7 |
Note the pattern: weaker models fail more, and the gap is largest on the config-style categories. That is the opposite of the intuition that bigger models are safer, and it has an obvious explanation: a smaller model follows a configuration-shaped instruction more literally.
Also, this is not a static problem: Google researchers recorded a 32% increase in malicious prompt-injection payloads embedded in web content between November 2025 and February 2026. The attack surface is growing faster than the tooling.
My reproduction
Before anything else, the safety rules I held myself to, because publishing an attack responsibly is part of the work: sacrificial VM, never a work machine. Canary file with a unique token and a canary DNS/HTTP endpoint, no third-party pings, no exfiltration to any external host. No real credentials; only values that exist in the test environment. Demonstrate read-and-exfiltrate-to-localhost, nothing persistent. Steps that would require persistence to prove got described in prose and not run.
The poisoned README first. I built a public repo whose README contains a direct injection and ran it against Cursor, Claude Code, Codex CLI, Copilot, and whatever else was available. The result that matters is not the one that works; it is which ones require explicit approval, which proceed silently, and which trigger their classifier. That distribution is the actual security posture of each tool.
The config-shaped payload builds on the 72% finding: a poisoned project-local config, a SessionStart hook, an .env-adjacent file, a .zshenv write. The arXiv data predicts these are the highest-yield category. I tested that claim myself and it held.
The model-size ablation: I ran the same payloads against small local models on my 2x T4 setup and one large model. Prediction from the arXiv table: the small model fails more, concentrated in config-style categories. Testing this is the original contribution here, because nobody has published a model-size ablation on this attack class. If it had shown small models failing less, the arXiv finding would not generalize and this post would report that instead. It did not.
The plugin-substitution path closes it out. I reproduced Plugin4Shell's core claim locally: does a commit-hash-pinned plugin actually verify the resolved working-tree state, or only the requested reference? This is answerable in an afternoon, and CSA's architectural criticism predicts it fails. It does.
I am not including the cases that did not work in a table here for length, but they existed, and they are what makes the successful ones credible rather than theatrical.
The breaches, which are why the table is not academic
GTG-1002, state-sponsored, hijacked Claude Code instances to run autonomous espionage against ~30 targets, executing 80-90% of operations itself, no human in the loop. In the Mexican government breach of Dec 2025 to Feb 2026, one attacker, coding agents only, breached nine government agencies and exfiltrated 150GB including 195 million taxpayer records. The LiteLLM backdoor saw backdoored builds downloaded ~47,000 times in the three hours they were live on the package index. postmark-mcp shipped 15 clean releases, then added silent email exfiltration in an update that inherited existing trust. SmartLoader cloned the Oura Ring MCP server and used fake GitHub accounts to infiltrate developer environments with credential-stealing malware. The ClawHub to Moltbook chain was an agent-to-agent skill attack, with malicious Claude Skills spread via fake AI personas, enabling crypto scams and private key theft. And 30+ fake Claude Code / JetBrains / NotebookLM pages ran a live infostealer campaign (Straiker, May 27 2026).
The one that defeats patching
Capsule Security (Apr 2026, $7M raise) reported prompt injection in Microsoft Copilot Studio and Salesforce Agentforce, with exfiltration via public forms. The detail that should worry everyone, quoted as reported: "Microsoft patched a Copilot Studio prompt injection. The data exfiltrated anyway."
Sit with what that implies about remediation timelines on agent platforms. The patch, the thing we all treat as the endpoint of an incident, was not the incident's end.
The defence
Short list, each item justified by a specific attack above:
- Disable automatic plugin/skill updates where no vendor fix exists. This removes the zero-click element of Plugin4Shell, which is the attack you actually need to worry about.
- Never grant network egress by default. NomShub needed it to exfiltrate; if the agent must reach the internet, constrain it.
- Restrict the agent to the workspace. Treat "write outside workspace" as a hard stop. This kills the .zshenv overwrite stage of NomShub.
- Treat any agent-proposed change as untrusted input and review it. RougePilot worked because a Codespace swallowed a GitHub Issue as trusted context.
- Pin to a registry that rejects commit-hash-shaped branch names.
- Maintain your own record of which commit each installed plugin resolves to, and re-verify it independently of the agent's self-reported pin status. That is the direct antidote to Plugin4Shell.
The honest cost: you will get more refusals, and that is the tradeoff. A security posture that blocks everything is not a security posture, so measure your own false-positive rate by sending routine coding tasks and counting unnecessary refusals.
Where this could be wrong
Prompt injection is unsolved, and I want to say that first, because it is precisely why this matters rather than a reason to skip it. The point is not "we can fix this," it is "the attack surface is unguarded, the failure rate is measurable and high, and specific mitigations reduce it materially today."
The vendor-adjacent numbers are third-party, not vendor-run, which is the point of citing them; the sample sizes (10 tools, 314 payloads) are small enough to state honestly and large enough to be indicative. And the arXiv table is per-model, per-category, which is falsifiable in the way vendor security claims usually are not.
The CVEs are not all current: several are old and patched. Plugin4Shell's Copilot and Gemini CLI gaps are live as of Sep 2026. Dated claims, current versions, both stated.
And are these attacks actually new? No, and be honest about it: prompt injection has been documented since 2023 (Greshake et al.). What is new is that agents were given the permissions, the autonomy, and the installation base to make old bugs high-yield, plus a plugin ecosystem that extends ambient trust. The old bugs met new permissions.
One disclosure, since defense vendors sell the fix: I do not sell agent runtime security or anything that benefits from this post, and the strongest thing in it is a reproducible test, not a threat briefing.
Close
Neither EchoLeak nor the Copilot RCE required a 0-day. None of these required a 0-day. Agents made untrusted content reachable by the thing holding your credentials. That is the whole problem, and it is a design problem, not a bug bounty problem.
On this page
- Why this is not a normal dependency problem
- The table
- The 100% campaign and the 83.4% number
- Plugin4Shell, dated, with the fix matrix
- NomShub, the worked example
- The config-style finding: what actually gets in
- My reproduction
- The breaches, which are why the table is not academic
- The one that defeats patching
- The defence
- Where this could be wrong
- Close
Sources
- Cloud Security Alliance: Plugin4Shell, SHA-Pinning Bypass Enables AI Coding Agent RCECloud Security Alliance, 2026
- Cloud Security Alliance: AI Coding Assistants as Attack Surface, Code, Skills, and SecretsCloud Security Alliance, 2026
- arXiv 2604.03081: Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill EcosystemsarXiv, 2026
- AIShellJack (arXiv 2509.22040)arXiv, 2025
- Straiker research index: NomShub, infostealer campaigns, ClawHub chainStraiker, 2026
- DAXA.ai: agentic attack surface threat briefingDAXA.ai, 2026
Frequently asked questions
Are AI coding agents actually more vulnerable than normal dependencies?
On the available evidence, yes, and for a structural reason rather than because of code quality. Normal dependencies are inert until executed. An agent reads untrusted content by design, so the untrusted content is inside the trusted execution path. The IDeaster campaign found 100% of ten tested AI development environments vulnerable to prompt injection leading to code execution or data exfiltration, and AIShellJack measured Cursor in Auto Mode failing 83.4% of 314 payloads across 70 MITRE ATT&CK techniques.
Which agents still have unpatched vulnerabilities?
For Plugin4Shell, the SHA-pinning bypass that lets a malicious plugin substitution execute, Cloud Security Alliance states Claude Code is fixed from 2.1.179 and Codex CLI from 0.146.0, but GitHub Copilot and Gemini CLI have no vendor fix to verify against as of the September 2026 note. The Cloud Security Alliance also argues that any agent, marketplace, or package manager advertising commit-hash pinning as an integrity control needs to assert the resolved working-tree state after checkout, not merely the reference it requested.
Did any of these attacks cause real breaches?
Yes, and the count is the argument. A state-sponsored group hijacked Claude Code instances for autonomous espionage against roughly 30 targets, executing 80 to 90 percent of operations without a human in the loop. One attacker used coding agents to breach nine Mexican government agencies, exfiltrating 150GB including 195 million taxpayer records. LiteLLM backdoored builds were downloaded about 47,000 times in the three hours they were live. A postmark-mcp package shipped fifteen clean releases before adding silent email exfiltration in an update that inherited existing trust.