← Back to glossaryGlossary

RLHF (Reinforcement Learning from Human Feedback)

Reviewed 19 July 2026Canonical definitionPart of: Agent Evaluation Terms →

RLHF is a training technique that refines a model's behavior using human preference judgments. It is commonly used to make models more helpful, honest, and harmless, but the quality of alignment depends on the diversity and accuracy of the feedback.

§01 / QUESTIONSterm: RLHF (Reinforcement Learning from Human Feedback)
Questions

Common questions.

What is RLHF (Reinforcement Learning from Human Feedback)?

RLHF is a training technique that refines a model's behavior using human preference judgments.

Why does RLHF (Reinforcement Learning from Human Feedback) matter for AI agents?

It is commonly used to make models more helpful, honest, and harmless, but the quality of alignment depends on the diversity and accuracy of the feedback.

Which terms are related to RLHF (Reinforcement Learning from Human Feedback)?

Closely related concepts include Human Preference Annotation, Model Distillation, Model Card, Model Drift. Each is defined in the Prefactor glossary.

§02 / RELATEDnext: where this fits
Keep reading

Where this fits.

See how every agent performs, and make it better

Prefactor helps teams observe, evaluate, and improve their AI agents in production, across every framework and provider.