Common questions.
What is Human Evaluation?
Human evaluation is the process of having people assess an AI agent's outputs for quality, accuracy, helpfulness, and safety.
How does Human Evaluation work?
It captures nuances that automated metrics miss and is essential for validating agents that handle subjective or high-stakes tasks.
Which terms are related to Human Evaluation?
Closely related concepts include Human Preference Annotation, Evaluation Harness (Agent), Ground Truth Evaluation, Quality Gate (AI). Each is defined in the Prefactor glossary.