Evaluation
LLM-as-judge
What is llm-as-judge?
LLM-as-judge uses a language model to grade the outputs of another model against a rubric. It scales subjective quality assessment far beyond what human review can cover, at the cost of introducing the judge’s own biases into the measurement.
Known biases are well documented: judges prefer longer responses, favour whichever output is presented first, and rate text stylistically similar to their own more highly. Randomising order and controlling for length are the minimum defences.
A judge should be validated before it is trusted. Have humans grade a sample, measure agreement, and treat the judge as an instrument whose calibration is known rather than as ground truth.
Related terms
Start with the free tier
A magic link creates your account, your tenant and your first API key. No card until you ask for compute.