Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917549976846336 |
|---|---|
| author | Kim, Yeongmin Shin, Donghyeok Na, Byeonghu Park, Minsang Kim, Richard Lee Moon, Il-Chul |
| author_facet | Kim, Yeongmin Shin, Donghyeok Na, Byeonghu Park, Minsang Kim, Richard Lee Moon, Il-Chul |
| contents | Diffusion models have demonstrated strong generative performance; however, generated samples often fail to fully align with human intent. This paper studies an efficient test-time scaling method for sampling from regions with higher human-aligned reward values. Existing methods for computing the expected future reward (EFR) face important limitations: backward rollout incurs prohibitively high sampling costs, while Tweedie-based approaches, including Sequential Monte Carlo and gradient guidance, suffer from bias and inherent sampling issues. We show that the EFR at any $\mathbf{x}_t$ can be computed using only marginal samples from a pre-trained diffusion model, enabling closed-form reward guidance without neural backpropagation. To further improve efficiency, we introduce a few-step lookahead sampling and an accurate solver that guides particles toward high-reward lookahead samples. We refer to this sampling scheme as LiDAR sampling. LiDAR achieves the same GenEval performance as the latest gradient guidance method for SDXL with a 9.5x speedup. We release the code at https://github.com/aailab-kaist/Diffusion-LiDAR-Sampling. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_03211 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models Kim, Yeongmin Shin, Donghyeok Na, Byeonghu Park, Minsang Kim, Richard Lee Moon, Il-Chul Machine Learning Artificial Intelligence Diffusion models have demonstrated strong generative performance; however, generated samples often fail to fully align with human intent. This paper studies an efficient test-time scaling method for sampling from regions with higher human-aligned reward values. Existing methods for computing the expected future reward (EFR) face important limitations: backward rollout incurs prohibitively high sampling costs, while Tweedie-based approaches, including Sequential Monte Carlo and gradient guidance, suffer from bias and inherent sampling issues. We show that the EFR at any $\mathbf{x}_t$ can be computed using only marginal samples from a pre-trained diffusion model, enabling closed-form reward guidance without neural backpropagation. To further improve efficiency, we introduce a few-step lookahead sampling and an accurate solver that guides particles toward high-reward lookahead samples. We refer to this sampling scheme as LiDAR sampling. LiDAR achieves the same GenEval performance as the latest gradient guidance method for SDXL with a 9.5x speedup. We release the code at https://github.com/aailab-kaist/Diffusion-LiDAR-Sampling. |
| title | Lookahead Sample Reward Guidance for Test-Time Scaling of Diffusion Models |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2602.03211 |