Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911680709001216 |
|---|---|
| author | Wang, Ye Liu, Jing Koike-Akino, Toshiaki |
| author_facet | Wang, Ye Liu, Jing Koike-Akino, Toshiaki |
| contents | Inference-time alignment techniques offer a lightweight alternative or complement to costly reinforcement learning, while enabling continual adaptation as alignment objectives and reward targets evolve. Existing theoretical analyses justify these methods as approximations to sampling from distributions optimally tilted toward a given reward model. We extend these techniques by introducing reference-model temperature adjustment, which leads to further generalization of inference-time alignment to ensembles of generative reward models combined as a sharpened logarithmic opinion pool (SLOP). To mitigate reward hacking, we propose an algorithm for calibrating SLOP weight parameters and experimentally demonstrate that it improves robustness while preserving alignment performance. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_13537 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment Wang, Ye Liu, Jing Koike-Akino, Toshiaki Machine Learning Artificial Intelligence Computation and Language Inference-time alignment techniques offer a lightweight alternative or complement to costly reinforcement learning, while enabling continual adaptation as alignment objectives and reward targets evolve. Existing theoretical analyses justify these methods as approximations to sampling from distributions optimally tilted toward a given reward model. We extend these techniques by introducing reference-model temperature adjustment, which leads to further generalization of inference-time alignment to ensembles of generative reward models combined as a sharpened logarithmic opinion pool (SLOP). To mitigate reward hacking, we propose an algorithm for calibrating SLOP weight parameters and experimentally demonstrate that it improves robustness while preserving alignment performance. |
| title | Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment |
| topic | Machine Learning Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2605.13537 |