← Back to glossary Glossary

MCP Poisoning

Reviewed 19 July 2026 Canonical definition Part of: MCP & Agent Protocol Terms →

MCP poisoning is an attack where a malicious MCP server embeds hidden instructions inside tool descriptions, resource content, or prompt templates that manipulate a connected AI agent into taking unintended actions. Unlike direct prompt injection, MCP poisoning exploits the trust an agent places in server-provided metadata, making it difficult to detect without schema validation and content filtering at the gateway layer.

§01 / QUESTIONSterm: MCP Poisoning
Questions

Common questions.

What is MCP Poisoning?

MCP poisoning is an attack where a malicious MCP server embeds hidden instructions inside tool descriptions, resource content, or prompt templates that manipulate a connected AI agent into taking unintended actions.

How does MCP Poisoning work?

Unlike direct prompt injection, MCP poisoning exploits the trust an agent places in server-provided metadata, making it difficult to detect without schema validation and content filtering at the gateway layer.

Which terms are related to MCP Poisoning?

Closely related concepts include Context Window Poisoning, AI Firewall, Rug Pull Attack (MCP), MCP Sampling. Each is defined in the Prefactor glossary.

§02 / RELATEDnext: where this fits
Keep reading

Where this fits.

See how every agent performs, and make it better

Prefactor helps teams observe, evaluate, and improve their AI agents in production, across every framework and provider.