AsymRnR: Video Diffusion Transformers Acceleration with Asymmetric Reduction and Restoration

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Wenhao, Tu, Rong-Cheng, Liao, Jingyi, Jin, Zhao, Tao, Dacheng
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916756031799296
author Sun, Wenhao
Tu, Rong-Cheng
Liao, Jingyi
Jin, Zhao
Tao, Dacheng
author_facet Sun, Wenhao
Tu, Rong-Cheng
Liao, Jingyi
Jin, Zhao
Tao, Dacheng
contents Diffusion Transformers (DiTs) have proven effective in generating high-quality videos but are hindered by high computational costs. Existing video DiT sampling acceleration methods often rely on costly fine-tuning or exhibit limited generalization capabilities. We propose Asymmetric Reduction and Restoration (AsymRnR), a training-free and model-agnostic method to accelerate video DiTs. It builds on the observation that redundancies of feature tokens in DiTs vary significantly across different model blocks, denoising steps, and feature types. Our AsymRnR asymmetrically reduces redundant tokens in the attention operation, achieving acceleration with negligible degradation in output quality and, in some cases, even improving it. We also tailored a reduction schedule to distribute the reduction across components adaptively. To further accelerate this process, we introduce a matching cache for more efficient reduction. Backed by theoretical foundations and extensive experimental validation, AsymRnR integrates into state-of-the-art video DiTs and offers substantial speedup.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11706
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AsymRnR: Video Diffusion Transformers Acceleration with Asymmetric Reduction and Restoration
Sun, Wenhao
Tu, Rong-Cheng
Liao, Jingyi
Jin, Zhao
Tao, Dacheng
Computer Vision and Pattern Recognition
Diffusion Transformers (DiTs) have proven effective in generating high-quality videos but are hindered by high computational costs. Existing video DiT sampling acceleration methods often rely on costly fine-tuning or exhibit limited generalization capabilities. We propose Asymmetric Reduction and Restoration (AsymRnR), a training-free and model-agnostic method to accelerate video DiTs. It builds on the observation that redundancies of feature tokens in DiTs vary significantly across different model blocks, denoising steps, and feature types. Our AsymRnR asymmetrically reduces redundant tokens in the attention operation, achieving acceleration with negligible degradation in output quality and, in some cases, even improving it. We also tailored a reduction schedule to distribute the reduction across components adaptively. To further accelerate this process, we introduce a matching cache for more efficient reduction. Backed by theoretical foundations and extensive experimental validation, AsymRnR integrates into state-of-the-art video DiTs and offers substantial speedup.
title AsymRnR: Video Diffusion Transformers Acceleration with Asymmetric Reduction and Restoration
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.11706