Common questions.
What is LLM-as-Judge?
LLM-as-judge is an evaluation technique where a separate language model, typically a capable model like GPT-4 or Claude, is used to score the outputs of an agent being evaluated.
Why does LLM-as-Judge matter for AI agents?
It enables scalable, automated assessment of subjective output qualities such as coherence, completeness, and tone that would otherwise require human annotators.
Which terms are related to LLM-as-Judge?
Closely related concepts include Ground Truth Evaluation, Hallucination Rate, Human Preference Annotation, Continuous Integration Agent. Each is defined in the Prefactor glossary.