首页 > AI前沿 > How to post-train on a surrogate: Envelope sampling mitigates reward hacking

How to post-train on a surrogate: Envelope sampling mitigates reward hacking

arXiv机器学习 2026-10-08 13:48 3 阅读 查看原文

Large language models (LLMs) are commonly post-trained against LLM judges and other cheap surrogates because the true reward, such as human preference, is too expensive to query at scale.

This practice often leads to reward hacking, where reinforcement learning against a miscalibrated surrogate leads to undesirable side effects.

In this work, we study a setting in which a small number $n$ of model outputs are annotated with ground-truth labels (e.g., from expert review) and used to recalibrate the LLM judge before optimizing against it.

Prior approaches to judge recalibration are costly or heuristic, and it is known that on-policy sampling fails when the surrogate is miscalibrated on a rare set of outputs.

In this work, we propose envelope sampling, a theoretically-grounded method for judge recalibration that seeks to minimize an upper bound on the regret of the post-trained model under the assumption that the human reward and re-calibrated reward lie in an $L^2$ ball around the judge.

We give practical algorithms to sample from the envelope by rejection or by fine-tuning against a modified reward.

And experiments on clinical note generation and on a controlled sycophancy task show that recalibrating on envelope samples mitigates reward hacking where recalibrating on base-model samples does not.