← Back to glossaryGlossary

LLM-as-Judge

Reviewed 19 July 2026Canonical definitionPart of: Agent Evaluation Terms →

LLM-as-judge is an evaluation technique where a separate language model, typically a capable model like GPT-4 or Claude, is used to score the outputs of an agent being evaluated. It enables scalable, automated assessment of subjective output qualities such as coherence, completeness, and tone that would otherwise require human annotators.

§01 / QUESTIONSterm: LLM-as-Judge
Questions

Common questions.

What is LLM-as-Judge?

LLM-as-judge is an evaluation technique where a separate language model, typically a capable model like GPT-4 or Claude, is used to score the outputs of an agent being evaluated.

Why does LLM-as-Judge matter for AI agents?

It enables scalable, automated assessment of subjective output qualities such as coherence, completeness, and tone that would otherwise require human annotators.

Which terms are related to LLM-as-Judge?

Closely related concepts include Ground Truth Evaluation, Hallucination Rate, Human Preference Annotation, Continuous Integration Agent. Each is defined in the Prefactor glossary.

§02 / RELATEDnext: where this fits
Keep reading

Where this fits.

See how every agent performs, and make it better

Prefactor helps teams observe, evaluate, and improve their AI agents in production, across every framework and provider.