As you crank up optimization pressure (KL distance from the base policy), the proxy reward the model is trained on keeps climbing — but the true reward peaks and then collapses. That gap is reward overoptimization, a.k.a. Goodhart's law. Drag the slider.
Same setup, different reward model. A pessimistic ensemble of reward models is far harder to hack — it pushes the achievable peak true-reward much higher before overoptimization sets in.