Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Yeongmin, Shin, Donghyeok, Na, Byeonghu, Park, Minsang, Kim, Richard Lee, Moon, Il-Chul
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917549976846336
author Kim, Yeongmin
Shin, Donghyeok
Na, Byeonghu
Park, Minsang
Kim, Richard Lee
Moon, Il-Chul
author_facet Kim, Yeongmin
Shin, Donghyeok
Na, Byeonghu
Park, Minsang
Kim, Richard Lee
Moon, Il-Chul
contents Diffusion models have demonstrated strong generative performance; however, generated samples often fail to fully align with human intent. This paper studies an efficient test-time scaling method for sampling from regions with higher human-aligned reward values. Existing methods for computing the expected future reward (EFR) face important limitations: backward rollout incurs prohibitively high sampling costs, while Tweedie-based approaches, including Sequential Monte Carlo and gradient guidance, suffer from bias and inherent sampling issues. We show that the EFR at any $\mathbf{x}_t$ can be computed using only marginal samples from a pre-trained diffusion model, enabling closed-form reward guidance without neural backpropagation. To further improve efficiency, we introduce a few-step lookahead sampling and an accurate solver that guides particles toward high-reward lookahead samples. We refer to this sampling scheme as LiDAR sampling. LiDAR achieves the same GenEval performance as the latest gradient guidance method for SDXL with a 9.5x speedup. We release the code at https://github.com/aailab-kaist/Diffusion-LiDAR-Sampling.
format Preprint
id arxiv_https___arxiv_org_abs_2602_03211
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models
Kim, Yeongmin
Shin, Donghyeok
Na, Byeonghu
Park, Minsang
Kim, Richard Lee
Moon, Il-Chul
Machine Learning
Artificial Intelligence
Diffusion models have demonstrated strong generative performance; however, generated samples often fail to fully align with human intent. This paper studies an efficient test-time scaling method for sampling from regions with higher human-aligned reward values. Existing methods for computing the expected future reward (EFR) face important limitations: backward rollout incurs prohibitively high sampling costs, while Tweedie-based approaches, including Sequential Monte Carlo and gradient guidance, suffer from bias and inherent sampling issues. We show that the EFR at any $\mathbf{x}_t$ can be computed using only marginal samples from a pre-trained diffusion model, enabling closed-form reward guidance without neural backpropagation. To further improve efficiency, we introduce a few-step lookahead sampling and an accurate solver that guides particles toward high-reward lookahead samples. We refer to this sampling scheme as LiDAR sampling. LiDAR achieves the same GenEval performance as the latest gradient guidance method for SDXL with a 9.5x speedup. We release the code at https://github.com/aailab-kaist/Diffusion-LiDAR-Sampling.
title Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.03211