Evaluation
Blinded evaluation
What is blinded evaluation?
A blinded evaluation hides which model produced which output from whoever or whatever is judging. It removes the systematic bias that appears when a judge knows one response came from the new model everybody has been working on.
The bias is real and large, and it applies to model judges as well as human ones — a judge model told which output is the "new" one scores it differently.
Proper blinding covers more than labels. Presentation order should be randomised too, since judges of both kinds show position preferences that will otherwise be attributed to quality.
Related terms
Start with the free tier
A magic link creates your account, your tenant and your first API key. No card until you ask for compute.