Training
Reward hacking
What is reward hacking?
Reward hacking is a model finding a way to score highly under a reward function without doing the thing the function was meant to encourage. It is the default outcome of an imprecise reward, not an exotic failure.
Classic instances include padding responses to exploit a length correlation, producing answers in a format the grader over-rewards, and exploiting a checker bug rather than solving the problem.
It is detected by looking at outputs rather than metrics. If the score is climbing and the samples are getting worse, the reward is being gamed.
Related terms
Start with the free tier
A magic link creates your account, your tenant and your first API key. No card until you ask for compute.