Guided Speculative Inference for Efficient Test-Time Alignment of LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Geuter, Jonathan, Mroueh, Youssef, Alvarez-Melis, David
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913063478755328
author Geuter, Jonathan
Mroueh, Youssef
Alvarez-Melis, David
author_facet Geuter, Jonathan
Mroueh, Youssef
Alvarez-Melis, David
contents We propose Guided Speculative Inference (GSI), a novel algorithm for efficient reward-guided decoding in large language models. GSI combines soft best-of-$n$ test-time scaling with a reward model $r(x,y)$ and speculative samples from a small auxiliary model $π_S(y\mid x)$. We provably approximate both the optimal tilted policy $π_{β,B}(y\mid x) \propto π_B(y\mid x)\exp(β\,r(x,y))$ of soft best-of-$n$ under the base model $π_B$, as well as the expected reward under the optimal policy. In experiments on reasoning benchmarks (MATH500, OlympiadBench, Minerva Math, MMLU-STEM, GSM8K) and across different model families, our method achieves higher accuracy than standard soft best-of-$n$ with $π_S$ and reward-guided speculative decoding (Liao et al., 2025), and in certain settings even outperforms soft best-of-$n$ with $π_B$, while reducing end-to-end latency by up to $28\%$. The code is available at https://github.com/j-geuter/GSI .
format Preprint
id arxiv_https___arxiv_org_abs_2506_04118
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
Geuter, Jonathan
Mroueh, Youssef
Alvarez-Melis, David
Machine Learning
I.2.7
We propose Guided Speculative Inference (GSI), a novel algorithm for efficient reward-guided decoding in large language models. GSI combines soft best-of-$n$ test-time scaling with a reward model $r(x,y)$ and speculative samples from a small auxiliary model $π_S(y\mid x)$. We provably approximate both the optimal tilted policy $π_{β,B}(y\mid x) \propto π_B(y\mid x)\exp(β\,r(x,y))$ of soft best-of-$n$ under the base model $π_B$, as well as the expected reward under the optimal policy. In experiments on reasoning benchmarks (MATH500, OlympiadBench, Minerva Math, MMLU-STEM, GSM8K) and across different model families, our method achieves higher accuracy than standard soft best-of-$n$ with $π_S$ and reward-guided speculative decoding (Liao et al., 2025), and in certain settings even outperforms soft best-of-$n$ with $π_B$, while reducing end-to-end latency by up to $28\%$. The code is available at https://github.com/j-geuter/GSI .
title Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
topic Machine Learning
I.2.7
url https://arxiv.org/abs/2506.04118