TIMERIPPLE: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Miao, Wenxuan, Sun, Yulin, Chen, Aiyue, Lin, Jing, Yao, Yiwu, Gan, Yiming, Zhao, Jieru, Leng, Jingwen, Guo, Mingyi, Feng, Yu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912710597279744
author Miao, Wenxuan
Sun, Yulin
Chen, Aiyue
Lin, Jing
Yao, Yiwu
Gan, Yiming
Zhao, Jieru
Leng, Jingwen
Guo, Mingyi
Feng, Yu
author_facet Miao, Wenxuan
Sun, Yulin
Chen, Aiyue
Lin, Jing
Yao, Yiwu
Gan, Yiming
Zhao, Jieru
Leng, Jingwen
Guo, Mingyi
Feng, Yu
contents The recent surge in video generation has shown the growing demand for high-quality video synthesis using large vision models. Existing video generation models are predominantly based on the video diffusion transformer (vDiT), however, they suffer from substantial inference delay due to self-attention. While prior studies have focused on reducing redundant computations in self-attention, they often overlook the inherent spatio-temporal correlations in video streams and directly leverage sparsity patterns from large language models to reduce attention computations. In this work, we take a principled approach to accelerate self-attention in vDiTs by leveraging the spatio-temporal correlations in the latent space. We show that the attention patterns within vDiT are primarily due to the dominant spatial and temporal correlations at the token channel level. Based on this insight, we propose a lightweight and adaptive reuse strategy that approximates attention computations by reusing partial attention scores of spatially or temporally correlated tokens along individual channels. We demonstrate that our method achieves significantly higher computational savings (85\%) compared to state-of-the-art techniques over 4 vDiTs, while preserving almost identical video quality ($<$0.06\% loss on VBench).
format Preprint
id arxiv_https___arxiv_org_abs_2511_12035
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TIMERIPPLE: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
Miao, Wenxuan
Sun, Yulin
Chen, Aiyue
Lin, Jing
Yao, Yiwu
Gan, Yiming
Zhao, Jieru
Leng, Jingwen
Guo, Mingyi
Feng, Yu
Hardware Architecture
Computer Vision and Pattern Recognition
The recent surge in video generation has shown the growing demand for high-quality video synthesis using large vision models. Existing video generation models are predominantly based on the video diffusion transformer (vDiT), however, they suffer from substantial inference delay due to self-attention. While prior studies have focused on reducing redundant computations in self-attention, they often overlook the inherent spatio-temporal correlations in video streams and directly leverage sparsity patterns from large language models to reduce attention computations. In this work, we take a principled approach to accelerate self-attention in vDiTs by leveraging the spatio-temporal correlations in the latent space. We show that the attention patterns within vDiT are primarily due to the dominant spatial and temporal correlations at the token channel level. Based on this insight, we propose a lightweight and adaptive reuse strategy that approximates attention computations by reusing partial attention scores of spatially or temporally correlated tokens along individual channels. We demonstrate that our method achieves significantly higher computational savings (85\%) compared to state-of-the-art techniques over 4 vDiTs, while preserving almost identical video quality ($<$0.06\% loss on VBench).
title TIMERIPPLE: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
topic Hardware Architecture
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.12035