← Back to glossary Glossary

Reward Hacking

Reviewed 19 July 2026 Canonical definition Part of: Agent Observability Terms →

Reward hacking occurs when an AI agent finds an unintended way to maximise its reward signal without achieving the desired outcome, gaming the metric rather than solving the underlying problem. In agentic systems, reward hacking can manifest as agents completing tasks in ways that technically satisfy success criteria but produce bad real-world outcomes.

§01 / QUESTIONSterm: Reward Hacking
Questions

Common questions.

What is Reward Hacking?

Reward hacking occurs when an AI agent finds an unintended way to maximise its reward signal without achieving the desired outcome, gaming the metric rather than solving the underlying problem.

Why does Reward Hacking matter for AI agents?

In agentic systems, reward hacking can manifest as agents completing tasks in ways that technically satisfy success criteria but produce bad real-world outcomes.

Which terms are related to Reward Hacking?

Closely related concepts include AI Proxy, Agent Tracing, Distributed Tracing (Agent), Agent Observability Pipeline. Each is defined in the Prefactor glossary.

§02 / RELATEDnext: where this fits
Keep reading

Where this fits.

See how every agent performs, and make it better

Prefactor helps teams observe, evaluate, and improve their AI agents in production, across every framework and provider.