onrup

Evaluation

LLM-as-judge

What is llm-as-judge?

LLM-as-judge uses a language model to grade the outputs of another model against a rubric. It scales subjective quality assessment far beyond what human review can cover, at the cost of introducing the judge’s own biases into the measurement.

Known biases are well documented: judges prefer longer responses, favour whichever output is presented first, and rate text stylistically similar to their own more highly. Randomising order and controlling for length are the minimum defences.

A judge should be validated before it is trusted. Have humans grade a sample, measure agreement, and treat the judge as an instrument whose calibration is known rather than as ground truth.

Related terms

All terms in the glossary →

Start with the free tier

A magic link creates your account, your tenant and your first API key. No card until you ask for compute.