Common questions.
What is RLHF (Reinforcement Learning from Human Feedback)?
RLHF is a training technique that refines a model's behavior using human preference judgments.
Why does RLHF (Reinforcement Learning from Human Feedback) matter for AI agents?
It is commonly used to make models more helpful, honest, and harmless, but the quality of alignment depends on the diversity and accuracy of the feedback.
Which terms are related to RLHF (Reinforcement Learning from Human Feedback)?
Closely related concepts include Human Preference Annotation, Model Distillation, Model Card, Model Drift. Each is defined in the Prefactor glossary.