Evaluation of Best-of-N Sampling Strategies for Language Model Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ichihara, Yuki, Jinnai, Yuu, Morimura, Tetsuro, Ariu, Kaito, Abe, Kenshi, Sakamoto, Mitsuki, Uchibe, Eiji
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916619290148864
author Ichihara, Yuki
Jinnai, Yuu
Morimura, Tetsuro
Ariu, Kaito
Abe, Kenshi
Sakamoto, Mitsuki
Uchibe, Eiji
author_facet Ichihara, Yuki
Jinnai, Yuu
Morimura, Tetsuro
Ariu, Kaito
Abe, Kenshi
Sakamoto, Mitsuki
Uchibe, Eiji
contents Best-of-N (BoN) sampling with a reward model has been shown to be an effective strategy for aligning Large Language Models (LLMs) with human preferences at the time of decoding. BoN sampling is susceptible to a problem known as reward hacking. Since the reward model is an imperfect proxy for the true objective, an excessive focus on optimizing its value can lead to a compromise of its performance on the true objective. Previous work proposes Regularized BoN sampling (RBoN), a BoN sampling with regularization to the objective, and shows that it outperforms BoN sampling so that it mitigates reward hacking and empirically (Jinnai et al., 2024). However, Jinnai et al. (2024) introduce RBoN based on a heuristic and they lack the analysis of why such regularization strategy improves the performance of BoN sampling. The aim of this study is to analyze the effect of BoN sampling on regularization strategies. Using the regularization strategies corresponds to robust optimization, which maximizes the worst case over a set of possible perturbations in the proxy reward. Although the theoretical guarantees are not directly applicable to RBoN, RBoN corresponds to a practical implementation. This paper proposes an extension of the RBoN framework, called Stochastic RBoN sampling (SRBoN), which is a theoretically guaranteed approach to worst-case RBoN in proxy reward. We then perform an empirical evaluation using the AlpacaFarm and Anthropic's hh-rlhf datasets to evaluate which factors of the regularization strategies contribute to the improvement of the true proxy reward. In addition, we also propose another simple RBoN method, the Sentence Length Regularized BoN, which has a better performance in the experiment as compared to the previous methods.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12668
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
Ichihara, Yuki
Jinnai, Yuu
Morimura, Tetsuro
Ariu, Kaito
Abe, Kenshi
Sakamoto, Mitsuki
Uchibe, Eiji
Computation and Language
Best-of-N (BoN) sampling with a reward model has been shown to be an effective strategy for aligning Large Language Models (LLMs) with human preferences at the time of decoding. BoN sampling is susceptible to a problem known as reward hacking. Since the reward model is an imperfect proxy for the true objective, an excessive focus on optimizing its value can lead to a compromise of its performance on the true objective. Previous work proposes Regularized BoN sampling (RBoN), a BoN sampling with regularization to the objective, and shows that it outperforms BoN sampling so that it mitigates reward hacking and empirically (Jinnai et al., 2024). However, Jinnai et al. (2024) introduce RBoN based on a heuristic and they lack the analysis of why such regularization strategy improves the performance of BoN sampling. The aim of this study is to analyze the effect of BoN sampling on regularization strategies. Using the regularization strategies corresponds to robust optimization, which maximizes the worst case over a set of possible perturbations in the proxy reward. Although the theoretical guarantees are not directly applicable to RBoN, RBoN corresponds to a practical implementation. This paper proposes an extension of the RBoN framework, called Stochastic RBoN sampling (SRBoN), which is a theoretically guaranteed approach to worst-case RBoN in proxy reward. We then perform an empirical evaluation using the AlpacaFarm and Anthropic's hh-rlhf datasets to evaluate which factors of the regularization strategies contribute to the improvement of the true proxy reward. In addition, we also propose another simple RBoN method, the Sentence Length Regularized BoN, which has a better performance in the experiment as compared to the previous methods.
title Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
topic Computation and Language
url https://arxiv.org/abs/2502.12668