Diffusion Alignment as Variational Expectation-Maximization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911491432644608 |
|---|---|
| author | Lee, Jaewoo Kim, Minsu Choi, Sanghyeok Song, Inhyuck Yun, Sujin Kang, Hyeongyu Shin, Woocheol Yun, Taeyoung Om, Kiyoung Park, Jinkyoo |
| author_facet | Lee, Jaewoo Kim, Minsu Choi, Sanghyeok Song, Inhyuck Yun, Sujin Kang, Hyeongyu Shin, Woocheol Yun, Taeyoung Om, Kiyoung Park, Jinkyoo |
| contents | Diffusion alignment aims to optimize diffusion models for the downstream objective. While existing methods based on reinforcement learning or direct backpropagation achieve considerable success in maximizing rewards, they often suffer from reward over-optimization and mode collapse. We introduce Diffusion Alignment as Variational Expectation-Maximization (DAV), a framework that formulates diffusion alignment as an iterative process alternating between two complementary phases: the E-step and the M-step. In the E-step, we employ test-time search to generate diverse and reward-aligned samples. In the M-step, we refine the diffusion model using samples discovered by the E-step. We demonstrate that DAV can optimize reward while preserving diversity for both continuous and discrete tasks: text-to-image synthesis and DNA sequence design. Our code is available at https://github.com/Jaewoopudding/dav. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_00502 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Diffusion Alignment as Variational Expectation-Maximization Lee, Jaewoo Kim, Minsu Choi, Sanghyeok Song, Inhyuck Yun, Sujin Kang, Hyeongyu Shin, Woocheol Yun, Taeyoung Om, Kiyoung Park, Jinkyoo Machine Learning Diffusion alignment aims to optimize diffusion models for the downstream objective. While existing methods based on reinforcement learning or direct backpropagation achieve considerable success in maximizing rewards, they often suffer from reward over-optimization and mode collapse. We introduce Diffusion Alignment as Variational Expectation-Maximization (DAV), a framework that formulates diffusion alignment as an iterative process alternating between two complementary phases: the E-step and the M-step. In the E-step, we employ test-time search to generate diverse and reward-aligned samples. In the M-step, we refine the diffusion model using samples discovered by the E-step. We demonstrate that DAV can optimize reward while preserving diversity for both continuous and discrete tasks: text-to-image synthesis and DNA sequence design. Our code is available at https://github.com/Jaewoopudding/dav. |
| title | Diffusion Alignment as Variational Expectation-Maximization |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2510.00502 |