Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Weilun, Yang, Chuanguang, Qin, Haotong, Li, Xiangqi, Wang, Yu, An, Zhulin, Huang, Libo, Diao, Boyu, Zhao, Zixiang, Xu, Yongjun, Magno, Michele
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918037224947712
author Feng, Weilun
Yang, Chuanguang
Qin, Haotong
Li, Xiangqi
Wang, Yu
An, Zhulin
Huang, Libo
Diao, Boyu
Zhao, Zixiang
Xu, Yongjun
Magno, Michele
author_facet Feng, Weilun
Yang, Chuanguang
Qin, Haotong
Li, Xiangqi
Wang, Yu
An, Zhulin
Huang, Libo
Diao, Boyu
Zhao, Zixiang
Xu, Yongjun
Magno, Michele
contents Diffusion transformers (DiT) have demonstrated exceptional performance in video generation. However, their large number of parameters and high computational complexity limit their deployment on edge devices. Quantization can reduce storage requirements and accelerate inference by lowering the bit-width of model parameters. Yet, existing quantization methods for image generation models do not generalize well to video generation tasks. We identify two primary challenges: the loss of information during quantization and the misalignment between optimization objectives and the unique requirements of video generation. To address these challenges, we present Q-VDiT, a quantization framework specifically designed for video DiT models. From the quantization perspective, we propose the Token-aware Quantization Estimator (TQE), which compensates for quantization errors in both the token and feature dimensions. From the optimization perspective, we introduce Temporal Maintenance Distillation (TMD), which preserves the spatiotemporal correlations between frames and enables the optimization of each frame with respect to the overall video context. Our W3A6 Q-VDiT achieves a scene consistency of 23.40, setting a new benchmark and outperforming current state-of-the-art quantization methods by 1.9$\times$. Code will be available at https://github.com/cantbebetter2/Q-VDiT.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22167
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers
Feng, Weilun
Yang, Chuanguang
Qin, Haotong
Li, Xiangqi
Wang, Yu
An, Zhulin
Huang, Libo
Diao, Boyu
Zhao, Zixiang
Xu, Yongjun
Magno, Michele
Computer Vision and Pattern Recognition
Diffusion transformers (DiT) have demonstrated exceptional performance in video generation. However, their large number of parameters and high computational complexity limit their deployment on edge devices. Quantization can reduce storage requirements and accelerate inference by lowering the bit-width of model parameters. Yet, existing quantization methods for image generation models do not generalize well to video generation tasks. We identify two primary challenges: the loss of information during quantization and the misalignment between optimization objectives and the unique requirements of video generation. To address these challenges, we present Q-VDiT, a quantization framework specifically designed for video DiT models. From the quantization perspective, we propose the Token-aware Quantization Estimator (TQE), which compensates for quantization errors in both the token and feature dimensions. From the optimization perspective, we introduce Temporal Maintenance Distillation (TMD), which preserves the spatiotemporal correlations between frames and enables the optimization of each frame with respect to the overall video context. Our W3A6 Q-VDiT achieves a scene consistency of 23.40, setting a new benchmark and outperforming current state-of-the-art quantization methods by 1.9$\times$. Code will be available at https://github.com/cantbebetter2/Q-VDiT.
title Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.22167