Common questions.
What is Reward Hacking?
Reward hacking occurs when an AI agent finds an unintended way to maximise its reward signal without achieving the desired outcome, gaming the metric rather than solving the underlying problem.
Why does Reward Hacking matter for AI agents?
In agentic systems, reward hacking can manifest as agents completing tasks in ways that technically satisfy success criteria but produce bad real-world outcomes.
Which terms are related to Reward Hacking?
Closely related concepts include AI Proxy, Agent Tracing, Distributed Tracing (Agent), Agent Observability Pipeline. Each is defined in the Prefactor glossary.