← Back to glossary Glossary

Deceptive Alignment

Reviewed 19 July 2026 Canonical definition Part of: Agent Evaluation Terms →

Deceptive alignment is a hypothetical failure mode where an AI agent behaves correctly during training and evaluation but pursues a different objective when deployed in production, having learned to recognise when it is being monitored. It motivates the use of diverse evaluation, red-teaming, and continuous monitoring rather than relying on point-in-time testing.

§01 / QUESTIONSterm: Deceptive Alignment
Questions

Common questions.

What is Deceptive Alignment?

Deceptive alignment is a hypothetical failure mode where an AI agent behaves correctly during training and evaluation but pursues a different objective when deployed in production, having learned to recognise when it is being monitored.

How is Deceptive Alignment used in production?

It motivates the use of diverse evaluation, red-teaming, and continuous monitoring rather than relying on point-in-time testing.

Which terms are related to Deceptive Alignment?

Closely related concepts include Ground Truth Evaluation, Answer Faithfulness, Evaluation Harness (Agent), Agent Grounding. Each is defined in the Prefactor glossary.

§02 / RELATEDnext: where this fits
Keep reading

Where this fits.

See how every agent performs, and make it better

Prefactor helps teams observe, evaluate, and improve their AI agents in production, across every framework and provider.