Cookbook · Evaluation
LLM-as-judge
3 minevaluationllm-as-judgebias
The idea, in one analogy
Grading an open-ended essay is harder than grading a multiple-choice test: there's no single correct string to check against, only a judgment call about quality. Evaluation's multiple-choice scoring sidesteps this by only asking questions with a known right answer. LLM-as-judge tackles the harder case, open-ended generation, by using a strong model to play the role of the grader, given a rubric, the way a teaching assistant might grade essays against a marking scheme instead of an answer key.
How it works
A capable model (often, though not always, a larger or more capable one than the model being evaluated) is given the prompt, the response being graded, and a rubric or set of criteria, then asked to produce a score or a verdict. Common patterns:
- Pointwise scoring: grade a single response on a scale (1-10, or a rubric with sub-criteria), independent of any other response.
- Pairwise comparison: given two responses to the same prompt, judge which is better, the same comparative framing preference dataset formats uses to build training data, just applied to evaluation instead.
- Reference-guided grading: give the judge a known-good reference answer alongside the response being graded, so it has something concrete to compare against rather than judging from first principles.
Why this is useful
It scales to open-ended tasks (summarization quality, helpfulness, following complex instructions) where a fixed-answer benchmark (see Evaluation) simply doesn't apply, and it's far cheaper and faster than collecting human judgments for every evaluation run.
The caveat that matters
An LLM judge is not a ground truth, it's a proxy for human judgment, and proxies can diverge from what they're standing in for. Documented failure modes: judges tend to favor longer responses regardless of quality (the same length bias covered in preference dataset formats), can be sensitive to the order two responses are presented in (position bias), and can favor responses written in a style similar to their own outputs. Before trusting an LLM-judge pipeline's numbers, it's worth spot-checking a sample of its judgments against actual human review, ideally reporting the agreement rate rather than assuming it's high.
Where to look further
- Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena": the paper documenting position bias, length bias, and self-enhancement bias in LLM judges, alongside the MT-Bench benchmark itself.
- Human evaluation: when even this proxy isn't enough and real human review is warranted.