TR-DQ: Time-Rotation Diffusion Quantization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shao, Yihua, Lin, Deyang, Zeng, Fanhu, Yan, Minxi, Zhang, Muyang, Chen, Siyu, Fan, Yuxuan, Yan, Ziyang, Wang, Haozhe, Guo, Jingcai, Wang, Yan, Qin, Haotong, Tang, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909532340355072
author Shao, Yihua
Lin, Deyang
Zeng, Fanhu
Yan, Minxi
Zhang, Muyang
Chen, Siyu
Fan, Yuxuan
Yan, Ziyang
Wang, Haozhe
Guo, Jingcai
Wang, Yan
Qin, Haotong
Tang, Hao
author_facet Shao, Yihua
Lin, Deyang
Zeng, Fanhu
Yan, Minxi
Zhang, Muyang
Chen, Siyu
Fan, Yuxuan
Yan, Ziyang
Wang, Haozhe
Guo, Jingcai
Wang, Yan
Qin, Haotong
Tang, Hao
contents Diffusion models have been widely adopted in image and video generation. However, their complex network architecture leads to high inference overhead for its generation process. Existing diffusion quantization methods primarily focus on the quantization of the model structure while ignoring the impact of time-steps variation during sampling. At the same time, most current approaches fail to account for significant activations that cannot be eliminated, resulting in substantial performance degradation after quantization. To address these issues, we propose Time-Rotation Diffusion Quantization (TR-DQ), a novel quantization method incorporating time-step and rotation-based optimization. TR-DQ first divides the sampling process based on time-steps and applies a rotation matrix to smooth activations and weights dynamically. For different time-steps, a dedicated hyperparameter is introduced for adaptive timing modeling, which enables dynamic quantization across different time steps. Additionally, we also explore the compression potential of Classifier-Free Guidance (CFG-wise) to establish a foundation for subsequent work. TR-DQ achieves state-of-the-art (SOTA) performance on image generation and video generation tasks and a 1.38-1.89x speedup and 1.97-2.58x memory reduction in inference compared to existing quantization methods.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06564
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TR-DQ: Time-Rotation Diffusion Quantization
Shao, Yihua
Lin, Deyang
Zeng, Fanhu
Yan, Minxi
Zhang, Muyang
Chen, Siyu
Fan, Yuxuan
Yan, Ziyang
Wang, Haozhe
Guo, Jingcai
Wang, Yan
Qin, Haotong
Tang, Hao
Computer Vision and Pattern Recognition
Diffusion models have been widely adopted in image and video generation. However, their complex network architecture leads to high inference overhead for its generation process. Existing diffusion quantization methods primarily focus on the quantization of the model structure while ignoring the impact of time-steps variation during sampling. At the same time, most current approaches fail to account for significant activations that cannot be eliminated, resulting in substantial performance degradation after quantization. To address these issues, we propose Time-Rotation Diffusion Quantization (TR-DQ), a novel quantization method incorporating time-step and rotation-based optimization. TR-DQ first divides the sampling process based on time-steps and applies a rotation matrix to smooth activations and weights dynamically. For different time-steps, a dedicated hyperparameter is introduced for adaptive timing modeling, which enables dynamic quantization across different time steps. Additionally, we also explore the compression potential of Classifier-Free Guidance (CFG-wise) to establish a foundation for subsequent work. TR-DQ achieves state-of-the-art (SOTA) performance on image generation and video generation tasks and a 1.38-1.89x speedup and 1.97-2.58x memory reduction in inference compared to existing quantization methods.
title TR-DQ: Time-Rotation Diffusion Quantization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.06564