Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Ye, Liu, Jing, Koike-Akino, Toshiaki
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911680709001216
author Wang, Ye
Liu, Jing
Koike-Akino, Toshiaki
author_facet Wang, Ye
Liu, Jing
Koike-Akino, Toshiaki
contents Inference-time alignment techniques offer a lightweight alternative or complement to costly reinforcement learning, while enabling continual adaptation as alignment objectives and reward targets evolve. Existing theoretical analyses justify these methods as approximations to sampling from distributions optimally tilted toward a given reward model. We extend these techniques by introducing reference-model temperature adjustment, which leads to further generalization of inference-time alignment to ensembles of generative reward models combined as a sharpened logarithmic opinion pool (SLOP). To mitigate reward hacking, we propose an algorithm for calibrating SLOP weight parameters and experimentally demonstrate that it improves robustness while preserving alignment performance.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13537
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
Wang, Ye
Liu, Jing
Koike-Akino, Toshiaki
Machine Learning
Artificial Intelligence
Computation and Language
Inference-time alignment techniques offer a lightweight alternative or complement to costly reinforcement learning, while enabling continual adaptation as alignment objectives and reward targets evolve. Existing theoretical analyses justify these methods as approximations to sampling from distributions optimally tilted toward a given reward model. We extend these techniques by introducing reference-model temperature adjustment, which leads to further generalization of inference-time alignment to ensembles of generative reward models combined as a sharpened logarithmic opinion pool (SLOP). To mitigate reward hacking, we propose an algorithm for calibrating SLOP weight parameters and experimentally demonstrate that it improves robustness while preserving alignment performance.
title Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2605.13537