Explicit Critic Guidance for Aligning Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liang, Zhengyang, Zhang, Qihang, Yang, Ceyuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914606715240448
author Liang, Zhengyang
Zhang, Qihang
Yang, Ceyuan
author_facet Liang, Zhengyang
Zhang, Qihang
Yang, Ceyuan
contents Online reinforcement learning is becoming increasingly important for aligning diffusion models with non-differentiable objectives. However, existing methods still face limitations in assigning fine-grained credit along denoising trajectories and in realizing stable value-based optimization. We propose a state-aligned latent actor-critic framework for diffusion post-training, in which the diffusion model serves as its own timestep-conditioned value function and predicts values directly on noisy latent states. This enables trajectory-level PPO training, supports stable actor-critic optimization with simple conditioning and value pretraining strategies, and naturally allows the learned critic to be reused for inference-time steering. We further extend the framework to multi-reward optimization, where joint training with complementary rewards helps alleviate reward hacking. Across both UNet- and DiT-based backbones, our method consistently outperforms prior group-relative RL and actor-critic baselines on single-reward and multi-reward benchmarks, while test-time steering provides additional gains in generation quality.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27736
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Explicit Critic Guidance for Aligning Diffusion Models
Liang, Zhengyang
Zhang, Qihang
Yang, Ceyuan
Machine Learning
Computer Vision and Pattern Recognition
Online reinforcement learning is becoming increasingly important for aligning diffusion models with non-differentiable objectives. However, existing methods still face limitations in assigning fine-grained credit along denoising trajectories and in realizing stable value-based optimization. We propose a state-aligned latent actor-critic framework for diffusion post-training, in which the diffusion model serves as its own timestep-conditioned value function and predicts values directly on noisy latent states. This enables trajectory-level PPO training, supports stable actor-critic optimization with simple conditioning and value pretraining strategies, and naturally allows the learned critic to be reused for inference-time steering. We further extend the framework to multi-reward optimization, where joint training with complementary rewards helps alleviate reward hacking. Across both UNet- and DiT-based backbones, our method consistently outperforms prior group-relative RL and actor-critic baselines on single-reward and multi-reward benchmarks, while test-time steering provides additional gains in generation quality.
title Explicit Critic Guidance for Aligning Diffusion Models
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.27736