← Back to glossary Glossary

Jailbreak (AI Agent)

Reviewed 19 July 2026 Canonical definition Part of: Runtime Control & Guardrail Terms →

A jailbreak is a technique used to bypass an AI agent's safety guardrails or governance constraints, typically through crafted prompts, role-play framings, or instruction injections that cause the agent to ignore its system prompt or policy rules. Unlike external attacks, jailbreaks often originate from end users attempting to expand what the agent will do. Runtime policy enforcement provides a defence layer that operates independently of the model's own safety training.

§01 / QUESTIONSterm: Jailbreak (AI Agent)
Questions

Common questions.

What is Jailbreak (AI Agent)?

A jailbreak is a technique used to bypass an AI agent's safety guardrails or governance constraints, typically through crafted prompts, role-play framings, or instruction injections that cause the agent to ignore its system prompt or policy rules.

How is Jailbreak (AI Agent) used in production?

Unlike external attacks, jailbreaks often originate from end users attempting to expand what the agent will do. Runtime policy enforcement provides a defence layer that operates independently of the model's own safety training.

Which terms are related to Jailbreak (AI Agent)?

Closely related concepts include Evasion Attack (AI), Guardrail Bypass, Capability Control (AI), Policy Enforcement Point. Each is defined in the Prefactor glossary.

§02 / RELATEDnext: where this fits
Keep reading

Where this fits.

See how every agent performs, and make it better

Prefactor helps teams observe, evaluate, and improve their AI agents in production, across every framework and provider.