Common questions.
What is Jailbreak (AI Agent)?
A jailbreak is a technique used to bypass an AI agent's safety guardrails or governance constraints, typically through crafted prompts, role-play framings, or instruction injections that cause the agent to ignore its system prompt or policy rules.
How is Jailbreak (AI Agent) used in production?
Unlike external attacks, jailbreaks often originate from end users attempting to expand what the agent will do. Runtime policy enforcement provides a defence layer that operates independently of the model's own safety training.
Which terms are related to Jailbreak (AI Agent)?
Closely related concepts include Evasion Attack (AI), Guardrail Bypass, Capability Control (AI), Policy Enforcement Point. Each is defined in the Prefactor glossary.