RLHF turns preference data into a training signal: humans (or proxies) rank answers, a reward model learns those rankings, then the policy model is optimized against it.

It is why assistants apologize, hedge, and follow chat norms that raw pretrained models lack.

Latent policy and labor stories increasingly ask who the raters are, what rubrics they use, and how preference data becomes a competitive asset.