Salvato in:
Dettagli Bibliografici
Autori principali: Kahatapitiya, Kumara, Liu, Haozhe, He, Sen, Liu, Ding, Jia, Menglin, Zhang, Chenyang, Ryoo, Michael S., Xie, Tian
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:https://arxiv.org/abs/2411.02397
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912108855164928
author Kahatapitiya, Kumara
Liu, Haozhe
He, Sen
Liu, Ding
Jia, Menglin
Zhang, Chenyang
Ryoo, Michael S.
Xie, Tian
author_facet Kahatapitiya, Kumara
Liu, Haozhe
He, Sen
Liu, Ding
Jia, Menglin
Zhang, Chenyang
Ryoo, Michael S.
Xie, Tian
contents Generating temporally-consistent high-fidelity videos can be computationally expensive, especially over longer temporal spans. More-recent Diffusion Transformers (DiTs) -- despite making significant headway in this context -- have only heightened such challenges as they rely on larger models and heavier attention mechanisms, resulting in slower inference speeds. In this paper, we introduce a training-free method to accelerate video DiTs, termed Adaptive Caching (AdaCache), which is motivated by the fact that "not all videos are created equal": meaning, some videos require fewer denoising steps to attain a reasonable quality than others. Building on this, we not only cache computations through the diffusion process, but also devise a caching schedule tailored to each video generation, maximizing the quality-latency trade-off. We further introduce a Motion Regularization (MoReg) scheme to utilize video information within AdaCache, essentially controlling the compute allocation based on motion content. Altogether, our plug-and-play contributions grant significant inference speedups (e.g. up to 4.7x on Open-Sora 720p - 2s video generation) without sacrificing the generation quality, across multiple video DiT baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02397
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Adaptive Caching for Faster Video Generation with Diffusion Transformers
Kahatapitiya, Kumara
Liu, Haozhe
He, Sen
Liu, Ding
Jia, Menglin
Zhang, Chenyang
Ryoo, Michael S.
Xie, Tian
Computer Vision and Pattern Recognition
Generating temporally-consistent high-fidelity videos can be computationally expensive, especially over longer temporal spans. More-recent Diffusion Transformers (DiTs) -- despite making significant headway in this context -- have only heightened such challenges as they rely on larger models and heavier attention mechanisms, resulting in slower inference speeds. In this paper, we introduce a training-free method to accelerate video DiTs, termed Adaptive Caching (AdaCache), which is motivated by the fact that "not all videos are created equal": meaning, some videos require fewer denoising steps to attain a reasonable quality than others. Building on this, we not only cache computations through the diffusion process, but also devise a caching schedule tailored to each video generation, maximizing the quality-latency trade-off. We further introduce a Motion Regularization (MoReg) scheme to utilize video information within AdaCache, essentially controlling the compute allocation based on motion content. Altogether, our plug-and-play contributions grant significant inference speedups (e.g. up to 4.7x on Open-Sora 720p - 2s video generation) without sacrificing the generation quality, across multiple video DiT baselines.
title Adaptive Caching for Faster Video Generation with Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.02397