TQ-DiT: Efficient Time-Aware Quantization for Diffusion Transformers

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Hwang, Younghye, Lee, Hyojin, Kang, Joonhyuk
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915141540380672
author Hwang, Younghye
Lee, Hyojin
Kang, Joonhyuk
author_facet Hwang, Younghye
Lee, Hyojin
Kang, Joonhyuk
contents Diffusion transformers (DiTs) combine transformer architectures with diffusion models. However, their computational complexity imposes significant limitations on real-time applications and sustainability of AI systems. In this study, we aim to enhance the computational efficiency through model quantization, which represents the weights and activation values with lower precision. Multi-region quantization (MRQ) is introduced to address the asymmetric distribution of network values in DiT blocks by allocating two scaling parameters to sub-regions. Additionally, time-grouping quantization (TGQ) is proposed to reduce quantization error caused by temporal variation in activations. The experimental results show that the proposed algorithm achieves performance comparable to the original full-precision model with only a 0.29 increase in FID at W8A8. Furthermore, it outperforms other baselines at W6A6, thereby confirming its suitability for low-bit quantization. These results highlight the potential of our method to enable efficient real-time generative models.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04056
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TQ-DiT: Efficient Time-Aware Quantization for Diffusion Transformers
Hwang, Younghye
Lee, Hyojin
Kang, Joonhyuk
Machine Learning
Signal Processing
Diffusion transformers (DiTs) combine transformer architectures with diffusion models. However, their computational complexity imposes significant limitations on real-time applications and sustainability of AI systems. In this study, we aim to enhance the computational efficiency through model quantization, which represents the weights and activation values with lower precision. Multi-region quantization (MRQ) is introduced to address the asymmetric distribution of network values in DiT blocks by allocating two scaling parameters to sub-regions. Additionally, time-grouping quantization (TGQ) is proposed to reduce quantization error caused by temporal variation in activations. The experimental results show that the proposed algorithm achieves performance comparable to the original full-precision model with only a 0.29 increase in FID at W8A8. Furthermore, it outperforms other baselines at W6A6, thereby confirming its suitability for low-bit quantization. These results highlight the potential of our method to enable efficient real-time generative models.
title TQ-DiT: Efficient Time-Aware Quantization for Diffusion Transformers
topic Machine Learning
Signal Processing
url https://arxiv.org/abs/2502.04056