← Back to glossary Glossary

Evasion Attack (AI)

Reviewed 19 July 2026 Canonical definition Part of: Runtime Control & Guardrail Terms →

An evasion attack crafts inputs that cause an AI agent's safety or policy filters to fail to detect policy violations, allowing harmful content, prompt injection, or unauthorised instructions to pass through governance controls. Evasion attacks test the robustness of agent guardrails and are a core component of AI red-teaming exercises.

§01 / QUESTIONSterm: Evasion Attack (AI)
Questions

Common questions.

What is Evasion Attack (AI)?

An evasion attack crafts inputs that cause an AI agent's safety or policy filters to fail to detect policy violations, allowing harmful content, prompt injection, or unauthorised instructions to pass through governance controls.

How does Evasion Attack (AI) work?

Evasion attacks test the robustness of agent guardrails and are a core component of AI red-teaming exercises.

Which terms are related to Evasion Attack (AI)?

Closely related concepts include Jailbreak (AI Agent), Capability Control (AI), Policy Enforcement Point, Agent Context Isolation. Each is defined in the Prefactor glossary.

§02 / RELATEDnext: where this fits
Keep reading

Where this fits.

See how every agent performs, and make it better

Prefactor helps teams observe, evaluate, and improve their AI agents in production, across every framework and provider.