Plugin4Shell: how a trick that dressed up code to look safe fooled the major AI models.
Updated: 4 days ago

In September 2026, cybersecurity company AIR Security uncovered Plugin4Shell, a vulnerability that exposed a weakness in the way major AI coding agents verify the plugins they run. The flaw affected Claude Code, OpenAI Codex, GitHub Copilot and Gemini CLI. It could allow a trusted plugin to be replaced with malicious code — without requiring the user to click or approve anything.
How does it work?
Plugins allow agents to perform actions and interact with other systems - something that is usually essential for writing code. If an attacker can get an agent to install malicious code pretending to be a useful plugin, they can cause it to execute harmful commands without ever influencing the behaviour of the agent, so it’s essential to make sure that any plugin your agent installs is safe.

Normally a system would do this by checking something called the git commit hash of the code they are installing. Git is a code management system that’s used almost universally, and it identifies any given version of any piece of code with the commit hash - a sort of fingerprint. It’s a 40 character sequence of letters and numbers, such as “'01f636aa7cc4115e2b77175c8b09d1ba31ccce77'”. If a marketplace has validated that a particular version of a piece code is safe it can publish the hash and when you install something you should be able to validate it by checking if the hash is on the list.
The attack works by redirecting an agent asking for a certified safe version of the code to an uncertified malicious version. This is possible because Git supports “references”, which is a way to rename a particular version of the code to something more human-readable. So instead of telling Git to “take me to '01f636aa7cc4115e2b77175c8b09d1ba31ccce77'” you can tell it to “take me to ‘version_where_I_fixed_a_bug’”, which is obviously very useful if you don’t want to remember a bunch of 40-character IDs. It’s a bit like entering a place name (Trafalgar Square) on Google Maps instead of a precise latitude and longitude (51°30'28.1"N 0°07'40.7"W).
What the attacker does is to take the commit hash that’s marked as safe, and then rename their malicious version to that hash. So when the agent tells Git “take me to '01f636aa7cc4115e2b77175c8b09d1ba31ccce77'”, it doesn’t go to that version, it goes to the malicious version. It’s as if someone had created a new street literally called “51°30'28.1"N 0°07'40.7"W”. Then when someone wants to find Trafalgar Square on Google Maps, they end up going to that street instead (and getting mugged). It’s possible to double-check the commit hash you end up on, but the exploit proved that the models weren’t doing this. They trusted an identifier without ensuring that it was resolved in the way the security control intended.
Can governance help with these types of vulnerabilities?
An AI agent is not just a model, it is better understood as a wider system:
Model + tools + plugins + data + permissions + external systems
Each component can change what the system can do and therefore what risks it introduces. Our earlier article, “What is Prompt Injection and how do I protect my company?”, examined what happens when an AI agent is manipulated by untrusted instructions. Plugin4Shell highlights a different risk: what happens when the agent runs code that the organisation mistakenly believes it has trusted? Both show why securing an AI system means looking beyond the model itself — to the tools, plugins and external components the agent relies on.
An organisation might say:
We approved this AI coding agent.
We reviewed its plugins.
We pinned plugins to known versions.
We have security controls around the agent.
But governance also requires answers to questions such as:
Which AI systems are actually in use?
What plugins, tools and external services do they depend on?
Who approved them and who is responsible?
What risks and controls were identified?
What has changed since the last assessment?
What evidence shows that controls are still working?
This is why AI governance needs to operate throughout the system lifecycle rather than being a one-time approval exercise. It provides a framework for establishing, implementing, maintaining and continually improving an AI Management System, including AI-related risks and responsibilities.
This doesn’t make an organisation immune to attacks like Plugin4Shell. But it forces you to look at every part of the system and think about what might go wrong. Even basic assumptions like “if a plugin version is verified it should be fine” can therefore be questioned, and weaknesses exposed. Some risks are inherently difficult to eliminate, particularly when systems depend on complex and changing components. Resilience therefore depends on knowing what you have, how it is connected, who is responsible for it, and where the relevant evidence is. If something does happen you have a track record of what assumptions you made, and will know which ones to think twice about in the future.
Think of it like a building in an earthquake zone. You cannot prevent an earthquake from happening, but knowing how the building was constructed—and having access to its structural plans—can make it easier to assess damage, coordinate a response and recover. AI systems are no different. When an unexpected incident occurs, visibility into the systems, components, risks, controls and responsibilities around an AI system can help an organisation respond and recover more effectively. This is where governance visibility becomes part of AI resilience.
If you need help setting up your AI management system we can help. Book a consultation with us here, or join the Beta for Vulpes, our purpose-built AI management system software.




Comments