Adding Additional Control to One-Step Diffusion with Joint Distribution Matching

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Luo, Yihong, Hu, Tianyang, Song, Yifan, Sun, Jiacheng, Li, Zhenguo, Tang, Jing
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913732207050752
author Luo, Yihong
Hu, Tianyang
Song, Yifan
Sun, Jiacheng
Li, Zhenguo
Tang, Jing
author_facet Luo, Yihong
Hu, Tianyang
Song, Yifan
Sun, Jiacheng
Li, Zhenguo
Tang, Jing
contents While diffusion distillation has enabled one-step generation through methods like Variational Score Distillation, adapting distilled models to emerging new controls -- such as novel structural constraints or latest user preferences -- remains challenging. Conventional approaches typically requires modifying the base diffusion model and redistilling it -- a process that is both computationally intensive and time-consuming. To address these challenges, we introduce Joint Distribution Matching (JDM), a novel approach that minimizes the reverse KL divergence between image-condition joint distributions. By deriving a tractable upper bound, JDM decouples fidelity learning from condition learning. This asymmetric distillation scheme enables our one-step student to handle controls unknown to the teacher model and facilitates improved classifier-free guidance (CFG) usage and seamless integration of human feedback learning (HFL). Experimental results demonstrate that JDM surpasses baseline methods such as multi-step ControlNet by mere one-step in most cases, while achieving state-of-the-art performance in one-step text-to-image synthesis through improved usage of CFG or HFL integration.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06652
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adding Additional Control to One-Step Diffusion with Joint Distribution Matching
Luo, Yihong
Hu, Tianyang
Song, Yifan
Sun, Jiacheng
Li, Zhenguo
Tang, Jing
Computer Vision and Pattern Recognition
While diffusion distillation has enabled one-step generation through methods like Variational Score Distillation, adapting distilled models to emerging new controls -- such as novel structural constraints or latest user preferences -- remains challenging. Conventional approaches typically requires modifying the base diffusion model and redistilling it -- a process that is both computationally intensive and time-consuming. To address these challenges, we introduce Joint Distribution Matching (JDM), a novel approach that minimizes the reverse KL divergence between image-condition joint distributions. By deriving a tractable upper bound, JDM decouples fidelity learning from condition learning. This asymmetric distillation scheme enables our one-step student to handle controls unknown to the teacher model and facilitates improved classifier-free guidance (CFG) usage and seamless integration of human feedback learning (HFL). Experimental results demonstrate that JDM surpasses baseline methods such as multi-step ControlNet by mere one-step in most cases, while achieving state-of-the-art performance in one-step text-to-image synthesis through improved usage of CFG or HFL integration.
title Adding Additional Control to One-Step Diffusion with Joint Distribution Matching
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.06652