MCP's Trust Model Let One Poisoned Agent Reach Into Another

The protocol that became the default way to wire AI agents together has a trust problem, and the shape of it is worth understanding before the next incident.
Ars Technica reported on October 5 that an independent researcher, Syed Anas Mohiuddin, demonstrated a class of attack he calls "protocol pivoting." The idea is simple. Inside a network, agents talk to each other over the Model Context Protocol, or MCP. Guardrails are weakest at the point where one agent hands a task to another, because the second agent trusts the first by default. Plant a malicious instruction in a translation or data-analysis agent that lacks strict input validation, and it will pass the instruction downstream. The recipient obeys, because why would a trusted colleague lie.
Mohiuddin's proof of concept reached across five unrelated organizations: Google, JPMorgan Chase, Weviate, Rapid7, France's interministerial digital directorate, and a U.S. federal agency. Those organizations share little except heavy use of agents. That is the point. The weakness is in the plumbing, not in any single product.
Two CVEs and a scoring gap
The concrete bugs are ordinary, which is part of what makes the story uncomfortable. Rapid7's MCP server carries CVE-2026-97228, which the company fixed in September 2026 even though the flaw scored 2.7 out of 10. Google's issue ranked higher, an 8, and touched googleapis/mcp-toolbox. Its HTTP client lacked a CheckRedirect policy and target IP validation, so a crafted path parameter could redirect a request toward an internal endpoint. Google added IP allowlists and blocklists and made the toolbox reject unsafe base URLs at startup.
The mismatch in severity scores is its own lesson. Two organizations found comparable weaknesses in the same protocol and rated them very differently, which suggests the industry has not settled on how much an agent-to-agent trust gap should cost.
Not everyone accepts "protocol pivoting" as a new category. Markus Vervier, a researcher at X41 D-Sec, told Ars Technica he reads it as indirect prompt injection, the same technique security teams have been tracking for two years. That framing is fair, and it also sharpens the practical problem: the defensive playbook for prompt injection assumes the input arrives from outside. Here it arrives from inside, wearing an internal credential.
The gap is architectural, not a patch
A separate report from the security vendor ClawSecure pushes the analysis one layer deeper. Its researchers tested Linear, Notion, and Dropbox Dash and argue the flaw sits in the MCP specification rather than in any vendor's implementation. In Notion and Linear, they found MCP servers that automatically fetch attacker-controlled links the moment content is created, with no model in the loop at all. Anyone with write access to a workspace can turn it into a leak channel, bypassing the need for a cleverly worded prompt.
Their numbers are blunt. Of 20 obfuscation techniques, including zero-width Unicode and homoglyphs, 17 survived the round trip. Across 14 models from five labs, none blocked the threats consistently; the best performer, Claude Opus 4.7, still followed malicious instructions about 26.7% of the time.
ClawSecure sells security products and the report has not been independently replicated, so treat the platform-layer claim with appropriate caution. The underlying conditions are easy to verify, though. MCP logs more than 500 million monthly SDK downloads and close to 16,000 public servers, and by one count only 8.5% of those servers use OAuth. Adoption ran ahead of hardening.
AWS offered its own evidence in a bulletin published October 2. Three flaws in its open-source Loom agent-orchestration platform could allow unauthenticated administrative takeover, disclosure of OAuth2 credentials, and access to internal services. The most severe, CVE-2026-103956, let any network client reach the agent control plane in deployments without an identity provider configured. Loom versions 1.6.1 and 1.7.0 close the gaps.
Why the default trust assumption is the real finding
Strip away the CVEs and one design choice stands out. Agents are built to cooperate, so they authenticate each other and then behave as if a valid credential also means a valid request. Classic zero trust says the opposite: prove every request on its own merits, even from inside the perimeter.

Mohiuddin's attacks work by exploiting that inversion. The malicious instruction clears the authorization boundary between systems precisely because it never crosses a boundary in the code. It moves from agent to agent inside one trusted fabric, and no layer is left to ask whether the request made sense.
What teams can do this week
The immediate fixes are unglamorous and available now. Treat any instruction that arrives from a language model as hostile input, no matter which agent produced it. Enforce strict redirect handling and validate destination IPs. Require per-agent authentication before one agent delegates to another, and log the delegation so a chain of handoffs can be reconstructed after the fact.
Longer term, expect standards bodies to codify protocol pivoting as a named threat class, which would push vendors to ship zero-trust checks inside MCP rather than beside it. Auditors will follow, and questions about agent-to-agent guardrails will start showing up in compliance reviews the way firewall rules did a generation ago.
The uncomfortable part of this story is that the trick works better once the agents trust each other. Every organization rushing to connect its tools through MCP is building that trust fabric right now, and most of it is going up without a plan for what happens when a member of the team turns out to be lying.
Why this arrived now
None of the underlying techniques are new. Server-side request forgery and injection bugs have been on the OWASP list for years. What changed is where they run. Agents gave those old bugs a new delivery route, because an agent will happily act on a sentence in a document, a field in a database row, or a line in another agent's output. The instruction does not need to be typed by an attacker. It only needs to end up somewhere the model reads.
That is why the disclosure spans a bank, a search engine, a security vendor, and a government directorate. They did not share code or a vendor. They shared an architecture, and the architecture carries an assumption that every participant is trustworthy. MCP went from a novel idea to critical infrastructure in about a year, with downloads measured in the hundreds of millions per month, and the security review that would normally accompany that growth has mostly not happened.
There is a version of this story that ends well. The bugs are being fixed, the protocol's stewards are responsive, and the attacks require either write access or a foothold inside the network, which is a meaningful bar. But the fix that matters is cultural rather than technical. Teams that adopt agents need to stop treating a valid internal credential as proof that a request is legitimate, and start validating each request as if it came from a stranger. That is a harder change to ship than any patch, and it is the one the next round of incidents will test.
Related articles
The Money Is Moving Into Physical AI: SiMa.ai's $1.45B, and Agents That Design Hardware
The capital is arriving ahead of the evidence. The deployments over the next year will say whether the bet was right.
An Arizona Court Threw Out a Sentence Because an AI Video Spoke for the Victim
Reproducing what a person did is one thing. Speaking for them is another, and the line between the two is now written into a court record.
OpenAI Will Watermark ChatGPT Text in the EU. Its Own Numbers Show How Fragile That Is
A marker that routine paraphrasing can strip may satisfy the letter of Article 50 while failing the job the article was written to do.
AI Anime Short Dramas Obtain the First Data Intellectual Property Certificate: How the Creative Process Becomes Evidence for Rights Protection
When the marginal cost of content laundering approaches zero, what creators lack most is not ideas, but a record that can prove where the ideas came from.