A General Framework for Inference-time Scaling and Steering of Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Singhal, Raghav, Horvitz, Zachary, Teehan, Ryan, Ren, Mengye, Yu, Zhou, McKeown, Kathleen, Ranganath, Rajesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913947567783936
author Singhal, Raghav
Horvitz, Zachary
Teehan, Ryan
Ren, Mengye
Yu, Zhou
McKeown, Kathleen
Ranganath, Rajesh
author_facet Singhal, Raghav
Horvitz, Zachary
Teehan, Ryan
Ren, Mengye
Yu, Zhou
McKeown, Kathleen
Ranganath, Rajesh
contents Diffusion models produce impressive results in modalities ranging from images and video to protein design and text. However, generating samples with user-specified properties remains a challenge. Recent research proposes fine-tuning models to maximize rewards that capture desired properties, but these methods require expensive training and are prone to mode collapse. In this work, we present Feynman-Kac (FK) steering, an inference-time framework for steering diffusion models with reward functions. FK steering works by sampling a system of multiple interacting diffusion processes, called particles, and resampling particles at intermediate steps based on scores computed using functions called potentials. Potentials are defined using rewards for intermediate states and are selected such that a high value indicates that the particle will yield a high-reward sample. We explore various choices of potentials, intermediate rewards, and samplers. We evaluate FK steering on text-to-image and text diffusion models. For steering text-to-image models with a human preference reward, we find that FK steering a 0.8B parameter model outperforms a 2.6B parameter fine-tuned model on prompt fidelity, with faster sampling and no training. For steering text diffusion models with rewards for text quality and specific text attributes, we find that FK steering generates lower perplexity, more linguistically acceptable outputs and enables gradient-free control of attributes like toxicity. Our results demonstrate that inference-time scaling and steering of diffusion models - even with off-the-shelf rewards - can provide significant sample quality gains and controllability benefits. Code is available at https://github.com/zacharyhorvitz/Fk-Diffusion-Steering .
format Preprint
id arxiv_https___arxiv_org_abs_2501_06848
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A General Framework for Inference-time Scaling and Steering of Diffusion Models
Singhal, Raghav
Horvitz, Zachary
Teehan, Ryan
Ren, Mengye
Yu, Zhou
McKeown, Kathleen
Ranganath, Rajesh
Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
Diffusion models produce impressive results in modalities ranging from images and video to protein design and text. However, generating samples with user-specified properties remains a challenge. Recent research proposes fine-tuning models to maximize rewards that capture desired properties, but these methods require expensive training and are prone to mode collapse. In this work, we present Feynman-Kac (FK) steering, an inference-time framework for steering diffusion models with reward functions. FK steering works by sampling a system of multiple interacting diffusion processes, called particles, and resampling particles at intermediate steps based on scores computed using functions called potentials. Potentials are defined using rewards for intermediate states and are selected such that a high value indicates that the particle will yield a high-reward sample. We explore various choices of potentials, intermediate rewards, and samplers. We evaluate FK steering on text-to-image and text diffusion models. For steering text-to-image models with a human preference reward, we find that FK steering a 0.8B parameter model outperforms a 2.6B parameter fine-tuned model on prompt fidelity, with faster sampling and no training. For steering text diffusion models with rewards for text quality and specific text attributes, we find that FK steering generates lower perplexity, more linguistically acceptable outputs and enables gradient-free control of attributes like toxicity. Our results demonstrate that inference-time scaling and steering of diffusion models - even with off-the-shelf rewards - can provide significant sample quality gains and controllability benefits. Code is available at https://github.com/zacharyhorvitz/Fk-Diffusion-Steering .
title A General Framework for Inference-time Scaling and Steering of Diffusion Models
topic Machine Learning
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.06848