Training
Reward function
What is reward function?
A reward function scores a model’s output during reinforcement learning. It is the complete specification of what the model will learn, and the model will optimise exactly what it measures rather than what was intended.
Programmatic checks — does this equal the expected answer, do these tests pass — are the most reliable form, because they are objective and cannot be talked around.
Every reward function has unintended optima. Watch for them explicitly, especially response length, which models reliably discover as a way to score better without being better.
Related terms
Start with the free tier
A magic link creates your account, your tenant and your first API key. No card until you ask for compute.