GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Yuancheng, Sehwag, Udari Madhushani, Koppel, Alec, Zhu, Sicheng, An, Bang, Huang, Furong, Ganesh, Sumitra
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913940531838976
author Xu, Yuancheng
Sehwag, Udari Madhushani
Koppel, Alec
Zhu, Sicheng
An, Bang
Huang, Furong
Ganesh, Sumitra
author_facet Xu, Yuancheng
Sehwag, Udari Madhushani
Koppel, Alec
Zhu, Sicheng
An, Bang
Huang, Furong
Ganesh, Sumitra
contents Large Language Models (LLMs) exhibit impressive capabilities but require careful alignment with human preferences. Traditional training-time methods finetune LLMs using human preference datasets but incur significant training costs and require repeated training to handle diverse user preferences. Test-time alignment methods address this by using reward models (RMs) to guide frozen LLMs without retraining. However, existing test-time approaches rely on trajectory-level RMs which are designed to evaluate complete responses, making them unsuitable for autoregressive text generation that requires computing next-token rewards from partial responses. To address this, we introduce GenARM, a test-time alignment approach that leverages the Autoregressive Reward Model--a novel reward parametrization designed to predict next-token rewards for efficient and effective autoregressive generation. Theoretically, we demonstrate that this parametrization can provably guide frozen LLMs toward any distribution achievable by traditional RMs within the KL-regularized reinforcement learning framework. Experimental results show that GenARM significantly outperforms prior test-time alignment baselines and matches the performance of training-time methods. Additionally, GenARM enables efficient weak-to-strong guidance, aligning larger LLMs with smaller RMs without the high costs of training larger models. Furthermore, GenARM supports multi-objective alignment, allowing real-time trade-offs between preference dimensions and catering to diverse user preferences without retraining. Our project page is available at: https://genarm.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2410_08193
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
Xu, Yuancheng
Sehwag, Udari Madhushani
Koppel, Alec
Zhu, Sicheng
An, Bang
Huang, Furong
Ganesh, Sumitra
Computation and Language
Large Language Models (LLMs) exhibit impressive capabilities but require careful alignment with human preferences. Traditional training-time methods finetune LLMs using human preference datasets but incur significant training costs and require repeated training to handle diverse user preferences. Test-time alignment methods address this by using reward models (RMs) to guide frozen LLMs without retraining. However, existing test-time approaches rely on trajectory-level RMs which are designed to evaluate complete responses, making them unsuitable for autoregressive text generation that requires computing next-token rewards from partial responses. To address this, we introduce GenARM, a test-time alignment approach that leverages the Autoregressive Reward Model--a novel reward parametrization designed to predict next-token rewards for efficient and effective autoregressive generation. Theoretically, we demonstrate that this parametrization can provably guide frozen LLMs toward any distribution achievable by traditional RMs within the KL-regularized reinforcement learning framework. Experimental results show that GenARM significantly outperforms prior test-time alignment baselines and matches the performance of training-time methods. Additionally, GenARM enables efficient weak-to-strong guidance, aligning larger LLMs with smaller RMs without the high costs of training larger models. Furthermore, GenARM supports multi-objective alignment, allowing real-time trade-offs between preference dimensions and catering to diverse user preferences without retraining. Our project page is available at: https://genarm.github.io.
title GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
topic Computation and Language
url https://arxiv.org/abs/2410.08193