SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ge, Xingtong, Zhang, Xin, Xu, Tongda, Zhang, Yi, Zhang, Xinjie, Wang, Yan, Zhang, Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910036168540160
author Ge, Xingtong
Zhang, Xin
Xu, Tongda
Zhang, Yi
Zhang, Xinjie
Wang, Yan
Zhang, Jun
author_facet Ge, Xingtong
Zhang, Xin
Xu, Tongda
Zhang, Yi
Zhang, Xinjie
Wang, Yan
Zhang, Jun
contents The Distribution Matching Distillation (DMD) has been successfully applied to text-to-image diffusion models such as Stable Diffusion (SD) 1.5. However, vanilla DMD suffers from convergence difficulties on large-scale flow-based text-to-image models, such as SD 3.5 and FLUX. In this paper, we first analyze the issues when applying vanilla DMD on large-scale models. Then, to overcome the scalability challenge, we propose implicit distribution alignment (IDA) to constrain the divergence between the generator and the fake distribution. Furthermore, we propose intra-segment guidance (ISG) to relocate the timestep denoising importance from the teacher model. With IDA alone, DMD converges for SD 3.5; employing both IDA and ISG, DMD converges for SD 3.5 and FLUX.1 dev. Together with a scaled VFM-based discriminator, our final model, dubbed \textbf{SenseFlow}, achieves superior performance in distillation for both diffusion based text-to-image models such as SDXL, and flow-matching models such as SD 3.5 Large and FLUX.1 dev. The source code is available at \href{https://github.com/XingtongGe/SenseFlow}{https://github.com/XingtongGe/SenseFlow}
format Preprint
id arxiv_https___arxiv_org_abs_2506_00523
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation
Ge, Xingtong
Zhang, Xin
Xu, Tongda
Zhang, Yi
Zhang, Xinjie
Wang, Yan
Zhang, Jun
Computer Vision and Pattern Recognition
The Distribution Matching Distillation (DMD) has been successfully applied to text-to-image diffusion models such as Stable Diffusion (SD) 1.5. However, vanilla DMD suffers from convergence difficulties on large-scale flow-based text-to-image models, such as SD 3.5 and FLUX. In this paper, we first analyze the issues when applying vanilla DMD on large-scale models. Then, to overcome the scalability challenge, we propose implicit distribution alignment (IDA) to constrain the divergence between the generator and the fake distribution. Furthermore, we propose intra-segment guidance (ISG) to relocate the timestep denoising importance from the teacher model. With IDA alone, DMD converges for SD 3.5; employing both IDA and ISG, DMD converges for SD 3.5 and FLUX.1 dev. Together with a scaled VFM-based discriminator, our final model, dubbed \textbf{SenseFlow}, achieves superior performance in distillation for both diffusion based text-to-image models such as SDXL, and flow-matching models such as SD 3.5 Large and FLUX.1 dev. The source code is available at \href{https://github.com/XingtongGe/SenseFlow}{https://github.com/XingtongGe/SenseFlow}
title SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.00523