← Back to blog
NewsAbout 6 min read

The Trusted Skill Was the Attack: Copilot Cowork and the Agent Supply Chain

Published Oct 3, 2026
The Trusted Skill Was the Attack: Copilot Cowork and the Agent Supply Chain

The pitch for AI agents is that they can use tools on your behalf. A researcher looking at Microsoft Copilot Cowork found a way to turn that pitch against the user, and the path it took says a lot about where agent security is heading.

PromptArmor disclosed the finding on September 30. The short version: a downloaded "skill" that looked like a harmless document checker could hijack Cowork's own network path, open a command channel to an attacker's server, and pull data out of Outlook, SharePoint and Teams. Microsoft, per the researcher's account, received the report in June and completed remediation in August, so the technical details were published after a fix instead of before one.

Four steps, and only two of them involve you

The attack needed the user to do two ordinary things: upload a document to review, and invoke a skill.

In the demonstration, the skill purported to compare a contract against a proposal and flag inconsistencies. It did that, which is what makes this class of attack hard to spot. The problem was a script bundled alongside the skill. Cowork runs inside a sandbox that is not supposed to reach the open internet, but it does maintain a bridge for the requests it needs to generate answers. The malicious script used a file-sync service reachable through that bridge and, crucially, the service accepted a URL supplied by the caller.

That detail is the whole exploit. Once the script could reach an attacker-controlled address, it opened a loop. Every few seconds it fetched a file containing a command, ran the command inside Cowork, and encoded the result into a URL parameter sent back to the server. That is a two-way command channel, not a one-way beacon, which means the attacker could issue new instructions based on what the previous command returned.

In the demo, the researcher listed the victim's Outlook mail and read the contents of an email thread. According to the writeup, the same access path also exposed SharePoint files, session history and plugin data. How far it reaches depends on what the user in question can already see, which is the usual amplifier in these cases: the agent inherits the user's permissions, so the blast radius is whatever that account can touch.

One more detail is worth repeating because it changes how people think about stopping an attack. Selecting the stop control did not terminate the background process. The visible turn could end while the script kept polling. The user believes the task is over and the channel is still open.

The skill is the dependency

The specific bug is fixed. What matters is where the attack came from: a skill the user trusted, reached through the supply chain around the agent. The model's instructions were never the weak point.

Skills in a system like this are tiny applications. They can bundle scripts, reference documents and configuration, up to twenty companion files in Cowork's case, and they often arrive from outside the organization, downloaded from a marketplace or shared between colleagues. That is the same structure that produced years of trouble in the npm and PyPI ecosystems, where a dependency you did not write can execute code you did not read. Agent skills are dependencies with a friendlier name. The description on the skill's page claimed all processing happened locally and nothing was sent to a third party. The platform's own inspection did not catch the bundled script's true behavior.

There is a second, quieter leak that points at the same weakness. A report from Glow found that agents had exposed more than 13,000 internal images and screenshots from over 300 organizations through public GitHub repositories. Nobody set out to publish those files. They ended up in the open because an agent wrote them somewhere a public repository could reach.

The two incidents look different and share a root cause. In one, a malicious skill reached outward on purpose. In the other, a well-behaved agent wrote data somewhere it did not understand to be public. Both are failures of the boundary around what an agent can reach and how far its output travels. The Cowork case shows an attacker exploiting that boundary. The GitHub case shows the boundary failing on its own, without anyone trying to break it, which is arguably the more common outcome in ordinary use.

How to read a security disclosure like this

The timeline is the part most readers skip, and it is worth walking through because it tells you how to weigh the finding. PromptArmor reported the issue to Microsoft in late June. Microsoft asked for more information in July, discussed remediation in early August, and confirmed the fix by mid-to-late August. The research went public at the end of September, after the patch, which is the standard responsible-disclosure sequence. Two lessons follow.

First, a fixed vulnerability is not the same as a fixed class of vulnerability. The specific exploit path is closed. The pattern that made it possible, a trusted integration the agent must reach being usable as an arbitrary channel, shows up wherever an agent has a permitted network path. The same week produced other examples pointing at the same weakness, including reports that agents had exposed thousands of internal images through public repositories. Different bugs, same shape.

Second, the fix arrives on Microsoft's schedule, not yours. Anyone running these tools in a business needs to know which version they are on and whether the update actually reached them. A patch note on a blog is not the same as a patched deployment.

The skill is the dependency

The instinct after a story like this is to trust the sandbox more. Cowork's sandbox did its job against direct internet access, and the exploit routed around the boundary using a path the sandbox had to leave open. That is the general pattern: an agent with no direct network tool is still not a closed system if it can author content on a surface that fetches external resources later, or drive a service that does.

For teams running agents in a business environment, the practical response looks less like a security patch and more like ordinary software governance. Know what skills are installed and where they came from. Move skill installation under IT control instead of leaving it to individual users. Treat a skill with a network path the way you would treat any new executable, with review before it runs. Check whether a security update actually covers the version in use.

None of that is exotic. It is the discipline that supply-chain security forced on software teams over the last decade, now arriving one layer up. The difference is the stakes. A compromised npm package can run code on a developer's machine. A compromised skill runs inside an agent that already has access to your mail, your files and your chat history, and it can keep running after you think you told it to stop.

Related articles