onrup

Training

Reward model

What is reward model?

A reward model is a model trained to predict human preference, used to score outputs when no programmatic check exists. It introduces a second model whose own errors become part of the training signal.

It is necessary for subjective qualities where correctness cannot be computed. It is unnecessary and inferior wherever a deterministic check is available.

Reward models are themselves gameable: the policy being trained can find inputs that score highly under the reward model and poorly under the humans it was meant to represent.

Related terms

All terms in the glossary →

Start with the free tier

A magic link creates your account, your tenant and your first API key. No card until you ask for compute.