TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fu, Minghao, Wang, Guo-Hua, Chen, Xiaohao, Chen, Qing-Guo, Xu, Zhao, Luo, Weihua, Zhang, Kaifu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909705055502336
author Fu, Minghao
Wang, Guo-Hua
Chen, Xiaohao
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
author_facet Fu, Minghao
Wang, Guo-Hua
Chen, Xiaohao
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
contents Recent advances in text-to-image synthesis largely benefit from sophisticated sampling strategies and classifier-free guidance (CFG) to ensure high-quality generation. However, CFG's reliance on two forward passes, especially when combined with intricate sampling algorithms, results in prohibitively high inference costs. To address this, we introduce TeEFusion (Text Embeddings Fusion), a novel and efficient distillation method that directly incorporates the guidance magnitude into the text embeddings and distills the teacher model's complex sampling strategy. By simply fusing conditional and unconditional text embeddings using linear operations, TeEFusion reconstructs the desired guidance without adding extra parameters, simultaneously enabling the student model to learn from the teacher's output produced via its sophisticated sampling approach. Extensive experiments on state-of-the-art models such as SD3 demonstrate that our method allows the student to closely mimic the teacher's performance with a far simpler and more efficient sampling strategy. Consequently, the student model achieves inference speeds up to 6$\times$ faster than the teacher model, while maintaining image quality at levels comparable to those obtained through the teacher's complex sampling approach. The code is publicly available at https://github.com/AIDC-AI/TeEFusion.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18192
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
Fu, Minghao
Wang, Guo-Hua
Chen, Xiaohao
Chen, Qing-Guo
Xu, Zhao
Luo, Weihua
Zhang, Kaifu
Computer Vision and Pattern Recognition
Recent advances in text-to-image synthesis largely benefit from sophisticated sampling strategies and classifier-free guidance (CFG) to ensure high-quality generation. However, CFG's reliance on two forward passes, especially when combined with intricate sampling algorithms, results in prohibitively high inference costs. To address this, we introduce TeEFusion (Text Embeddings Fusion), a novel and efficient distillation method that directly incorporates the guidance magnitude into the text embeddings and distills the teacher model's complex sampling strategy. By simply fusing conditional and unconditional text embeddings using linear operations, TeEFusion reconstructs the desired guidance without adding extra parameters, simultaneously enabling the student model to learn from the teacher's output produced via its sophisticated sampling approach. Extensive experiments on state-of-the-art models such as SD3 demonstrate that our method allows the student to closely mimic the teacher's performance with a far simpler and more efficient sampling strategy. Consequently, the student model achieves inference speeds up to 6$\times$ faster than the teacher model, while maintaining image quality at levels comparable to those obtained through the teacher's complex sampling approach. The code is publicly available at https://github.com/AIDC-AI/TeEFusion.
title TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.18192