Common questions.
What is Backdoor Attack (AI)?
A backdoor attack embeds hidden behaviour into an AI model during training, causing the model to behave normally under most inputs but to produce attacker-controlled outputs when a specific trigger pattern is present.
How is Backdoor Attack (AI) used in production?
Backdoors in foundation models used by AI agents can be extremely difficult to detect without extensive red-teaming.
Which terms are related to Backdoor Attack (AI)?
Closely related concepts include Pre-Training, Delegation Control, Sidecar Pattern (Agent), Episodic Memory (Agent). Each is defined in the Prefactor glossary.