← Back to glossary Glossary

Backdoor Attack (AI)

Reviewed 19 July 2026 Canonical definition Part of: Agent Orchestration & Workflow Terms →

A backdoor attack embeds hidden behaviour into an AI model during training, causing the model to behave normally under most inputs but to produce attacker-controlled outputs when a specific trigger pattern is present. Backdoors in foundation models used by AI agents can be extremely difficult to detect without extensive red-teaming.

§01 / QUESTIONSterm: Backdoor Attack (AI)
Questions

Common questions.

What is Backdoor Attack (AI)?

A backdoor attack embeds hidden behaviour into an AI model during training, causing the model to behave normally under most inputs but to produce attacker-controlled outputs when a specific trigger pattern is present.

How is Backdoor Attack (AI) used in production?

Backdoors in foundation models used by AI agents can be extremely difficult to detect without extensive red-teaming.

Which terms are related to Backdoor Attack (AI)?

Closely related concepts include Pre-Training, Delegation Control, Sidecar Pattern (Agent), Episodic Memory (Agent). Each is defined in the Prefactor glossary.

§02 / RELATEDnext: where this fits
Keep reading

Where this fits.

See how every agent performs, and make it better

Prefactor helps teams observe, evaluate, and improve their AI agents in production, across every framework and provider.