Common questions.
What is Deceptive Alignment?
Deceptive alignment is a hypothetical failure mode where an AI agent behaves correctly during training and evaluation but pursues a different objective when deployed in production, having learned to recognise when it is being monitored.
How is Deceptive Alignment used in production?
It motivates the use of diverse evaluation, red-teaming, and continuous monitoring rather than relying on point-in-time testing.
Which terms are related to Deceptive Alignment?
Closely related concepts include Continuous Integration Agent, Ground Truth Evaluation, Answer Faithfulness, Evaluation Harness (Agent). Each is defined in the Prefactor glossary.