One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Senmao, Wang, Lei, Wang, Kai, Liu, Tao, Xie, Jiehang, van de Weijer, Joost, Khan, Fahad Shahbaz, Yang, Shiqi, Wang, Yaxing, Yang, Jian
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910972542713856
author Li, Senmao
Wang, Lei
Wang, Kai
Liu, Tao
Xie, Jiehang
van de Weijer, Joost
Khan, Fahad Shahbaz
Yang, Shiqi
Wang, Yaxing
Yang, Jian
author_facet Li, Senmao
Wang, Lei
Wang, Kai
Liu, Tao
Xie, Jiehang
van de Weijer, Joost
Khan, Fahad Shahbaz
Yang, Shiqi
Wang, Yaxing
Yang, Jian
contents Text-to-Image (T2I) diffusion models have made remarkable advancements in generative modeling; however, they face a trade-off between inference speed and image quality, posing challenges for efficient deployment. Existing distilled T2I models can generate high-fidelity images with fewer sampling steps, but often struggle with diversity and quality, especially in one-step models. From our analysis, we observe redundant computations in the UNet encoders. Our findings suggest that, for T2I diffusion models, decoders are more adept at capturing richer and more explicit semantic information, while encoders can be effectively shared across decoders from diverse time steps. Based on these observations, we introduce the first Time-independent Unified Encoder TiUE for the student model UNet architecture, which is a loop-free image generation approach for distilling T2I diffusion models. Using a one-pass scheme, TiUE shares encoder features across multiple decoder time steps, enabling parallel sampling and significantly reducing inference time complexity. In addition, we incorporate a KL divergence term to regularize noise prediction, which enhances the perceptual realism and diversity of the generated images. Experimental results demonstrate that TiUE outperforms state-of-the-art methods, including LCM, SD-Turbo, and SwiftBrushv2, producing more diverse and realistic results while maintaining the computational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2505_21960
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models
Li, Senmao
Wang, Lei
Wang, Kai
Liu, Tao
Xie, Jiehang
van de Weijer, Joost
Khan, Fahad Shahbaz
Yang, Shiqi
Wang, Yaxing
Yang, Jian
Computer Vision and Pattern Recognition
Text-to-Image (T2I) diffusion models have made remarkable advancements in generative modeling; however, they face a trade-off between inference speed and image quality, posing challenges for efficient deployment. Existing distilled T2I models can generate high-fidelity images with fewer sampling steps, but often struggle with diversity and quality, especially in one-step models. From our analysis, we observe redundant computations in the UNet encoders. Our findings suggest that, for T2I diffusion models, decoders are more adept at capturing richer and more explicit semantic information, while encoders can be effectively shared across decoders from diverse time steps. Based on these observations, we introduce the first Time-independent Unified Encoder TiUE for the student model UNet architecture, which is a loop-free image generation approach for distilling T2I diffusion models. Using a one-pass scheme, TiUE shares encoder features across multiple decoder time steps, enabling parallel sampling and significantly reducing inference time complexity. In addition, we incorporate a KL divergence term to regularize noise prediction, which enhances the perceptual realism and diversity of the generated images. Experimental results demonstrate that TiUE outperforms state-of-the-art methods, including LCM, SD-Turbo, and SwiftBrushv2, producing more diverse and realistic results while maintaining the computational efficiency.
title One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.21960