Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Bolian, Wu, Yanran, Luo, Xinyu, Zhang, Ruqi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911170866184192
author Li, Bolian
Wu, Yanran
Luo, Xinyu
Zhang, Ruqi
author_facet Li, Bolian
Wu, Yanran
Luo, Xinyu
Zhang, Ruqi
contents Aligning large language models (LLMs) with human preferences has become a critical step in their development. Recent research has increasingly focused on test-time alignment, where additional compute is allocated during inference to enhance LLM safety and reasoning capabilities. However, these test-time alignment techniques often incur substantial inference costs, limiting their practical application. We are inspired by the speculative sampling acceleration, which leverages a small draft model to efficiently predict future tokens, to address the efficiency bottleneck of test-time alignment. We introduce the reward-shifted speculative sampling (SSS) algorithm, in which the draft model is aligned with human preferences, while the target model remains unchanged. We theoretically demonstrate that the distributional shift between the aligned draft model and the unaligned target model can be exploited to recover the RLHF optimal solution without actually obtaining it, by modifying the acceptance criterion and bonus token distribution. Our algorithm achieves superior gold reward scores at a significantly reduced inference cost in test-time weak-to-strong alignment experiments, thereby validating both its effectiveness and efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15044
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
Li, Bolian
Wu, Yanran
Luo, Xinyu
Zhang, Ruqi
Computation and Language
Aligning large language models (LLMs) with human preferences has become a critical step in their development. Recent research has increasingly focused on test-time alignment, where additional compute is allocated during inference to enhance LLM safety and reasoning capabilities. However, these test-time alignment techniques often incur substantial inference costs, limiting their practical application. We are inspired by the speculative sampling acceleration, which leverages a small draft model to efficiently predict future tokens, to address the efficiency bottleneck of test-time alignment. We introduce the reward-shifted speculative sampling (SSS) algorithm, in which the draft model is aligned with human preferences, while the target model remains unchanged. We theoretically demonstrate that the distributional shift between the aligned draft model and the unaligned target model can be exploited to recover the RLHF optimal solution without actually obtaining it, by modifying the acceptance criterion and bonus token distribution. Our algorithm achieves superior gold reward scores at a significantly reduced inference cost in test-time weak-to-strong alignment experiments, thereby validating both its effectiveness and efficiency.
title Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
topic Computation and Language
url https://arxiv.org/abs/2508.15044