Text Embedding Knows How to Quantize Text-Guided Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Hongjae, Son, Myungjun, Kang, Dongjea, Jung, Seung-Won
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912517988548608
author Lee, Hongjae
Son, Myungjun
Kang, Dongjea
Jung, Seung-Won
author_facet Lee, Hongjae
Son, Myungjun
Kang, Dongjea
Jung, Seung-Won
contents Despite the success of diffusion models in image generation tasks such as text-to-image, the enormous computational complexity of diffusion models limits their use in resource-constrained environments. To address this, network quantization has emerged as a promising solution for designing efficient diffusion models. However, existing diffusion model quantization methods do not consider input conditions, such as text prompts, as an essential source of information for quantization. In this paper, we propose a novel quantization method dubbed Quantization of Language-to-Image diffusion models using text Prompts (QLIP). QLIP leverages text prompts to guide the selection of bit precision for every layer at each time step. In addition, QLIP can be seamlessly integrated into existing quantization methods to enhance quantization efficiency. Our extensive experiments demonstrate the effectiveness of QLIP in reducing computational complexity and improving the quality of the generated images across various datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10340
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text Embedding Knows How to Quantize Text-Guided Diffusion Models
Lee, Hongjae
Son, Myungjun
Kang, Dongjea
Jung, Seung-Won
Computer Vision and Pattern Recognition
Despite the success of diffusion models in image generation tasks such as text-to-image, the enormous computational complexity of diffusion models limits their use in resource-constrained environments. To address this, network quantization has emerged as a promising solution for designing efficient diffusion models. However, existing diffusion model quantization methods do not consider input conditions, such as text prompts, as an essential source of information for quantization. In this paper, we propose a novel quantization method dubbed Quantization of Language-to-Image diffusion models using text Prompts (QLIP). QLIP leverages text prompts to guide the selection of bit precision for every layer at each time step. In addition, QLIP can be seamlessly integrated into existing quantization methods to enhance quantization efficiency. Our extensive experiments demonstrate the effectiveness of QLIP in reducing computational complexity and improving the quality of the generated images across various datasets.
title Text Embedding Knows How to Quantize Text-Guided Diffusion Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.10340