← Back to glossaryGlossary

Corrigibility

Reviewed 19 July 2026Canonical definitionPart of: AI Agent Fundamentals →

Corrigibility is the property of an AI agent that makes it responsive to correction, shutdown, and modification by its operators, without resisting, circumventing, or manipulating humans in order to preserve its current goals. Ensuring AI agents remain corrigible is a core AI safety objective, especially as agents become more capable of taking autonomous actions.

§01 / QUESTIONSterm: Corrigibility
Questions

Common questions.

What is Corrigibility?

Corrigibility is the property of an AI agent that makes it responsive to correction, shutdown, and modification by its operators, without resisting, circumventing, or manipulating humans in order to preserve its current goals.

How does Corrigibility work?

Ensuring AI agents remain corrigible is a core AI safety objective, especially as agents become more capable of taking autonomous actions.

Which terms are related to Corrigibility?

Closely related concepts include Event Sourcing (Agent), Value Alignment, AI Singularity, Artificial General Intelligence (AGI). Each is defined in the Prefactor glossary.

§02 / RELATEDnext: where this fits
Keep reading

Where this fits.

See how every agent performs, and make it better

Prefactor helps teams observe, evaluate, and improve their AI agents in production, across every framework and provider.