Diffusion Alignment as Variational Expectation-Maximization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Jaewoo, Kim, Minsu, Choi, Sanghyeok, Song, Inhyuck, Yun, Sujin, Kang, Hyeongyu, Shin, Woocheol, Yun, Taeyoung, Om, Kiyoung, Park, Jinkyoo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911491432644608
author Lee, Jaewoo
Kim, Minsu
Choi, Sanghyeok
Song, Inhyuck
Yun, Sujin
Kang, Hyeongyu
Shin, Woocheol
Yun, Taeyoung
Om, Kiyoung
Park, Jinkyoo
author_facet Lee, Jaewoo
Kim, Minsu
Choi, Sanghyeok
Song, Inhyuck
Yun, Sujin
Kang, Hyeongyu
Shin, Woocheol
Yun, Taeyoung
Om, Kiyoung
Park, Jinkyoo
contents Diffusion alignment aims to optimize diffusion models for the downstream objective. While existing methods based on reinforcement learning or direct backpropagation achieve considerable success in maximizing rewards, they often suffer from reward over-optimization and mode collapse. We introduce Diffusion Alignment as Variational Expectation-Maximization (DAV), a framework that formulates diffusion alignment as an iterative process alternating between two complementary phases: the E-step and the M-step. In the E-step, we employ test-time search to generate diverse and reward-aligned samples. In the M-step, we refine the diffusion model using samples discovered by the E-step. We demonstrate that DAV can optimize reward while preserving diversity for both continuous and discrete tasks: text-to-image synthesis and DNA sequence design. Our code is available at https://github.com/Jaewoopudding/dav.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00502
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Diffusion Alignment as Variational Expectation-Maximization
Lee, Jaewoo
Kim, Minsu
Choi, Sanghyeok
Song, Inhyuck
Yun, Sujin
Kang, Hyeongyu
Shin, Woocheol
Yun, Taeyoung
Om, Kiyoung
Park, Jinkyoo
Machine Learning
Diffusion alignment aims to optimize diffusion models for the downstream objective. While existing methods based on reinforcement learning or direct backpropagation achieve considerable success in maximizing rewards, they often suffer from reward over-optimization and mode collapse. We introduce Diffusion Alignment as Variational Expectation-Maximization (DAV), a framework that formulates diffusion alignment as an iterative process alternating between two complementary phases: the E-step and the M-step. In the E-step, we employ test-time search to generate diverse and reward-aligned samples. In the M-step, we refine the diffusion model using samples discovered by the E-step. We demonstrate that DAV can optimize reward while preserving diversity for both continuous and discrete tasks: text-to-image synthesis and DNA sequence design. Our code is available at https://github.com/Jaewoopudding/dav.
title Diffusion Alignment as Variational Expectation-Maximization
topic Machine Learning
url https://arxiv.org/abs/2510.00502