Training-Free Efficient Video Generation via Dynamic Token Carving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuechen, Xing, Jinbo, Xia, Bin, Liu, Shaoteng, Peng, Bohao, Tao, Xin, Wan, Pengfei, Lo, Eric, Jia, Jiaya
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911279340322816
author Zhang, Yuechen
Xing, Jinbo
Xia, Bin
Liu, Shaoteng
Peng, Bohao
Tao, Xin
Wan, Pengfei
Lo, Eric
Jia, Jiaya
author_facet Zhang, Yuechen
Xing, Jinbo
Xia, Bin
Liu, Shaoteng
Peng, Bohao
Tao, Xin
Wan, Pengfei
Lo, Eric
Jia, Jiaya
contents Despite the remarkable generation quality of video Diffusion Transformer (DiT) models, their practical deployment is severely hindered by extensive computational requirements. This inefficiency stems from two key challenges: the quadratic complexity of self-attention with respect to token length and the multi-step nature of diffusion models. To address these limitations, we present Jenga, a novel inference pipeline that combines dynamic attention carving with progressive resolution generation. Our approach leverages two key insights: (1) early denoising steps do not require high-resolution latents, and (2) later steps do not require dense attention. Jenga introduces a block-wise attention mechanism that dynamically selects relevant token interactions using 3D space-filling curves, alongside a progressive resolution strategy that gradually increases latent resolution during generation. Experimental results demonstrate that Jenga achieves substantial speedups across multiple state-of-the-art video diffusion models while maintaining comparable generation quality (8.83$\times$ speedup with 0.01\% performance drop on VBench). As a plug-and-play solution, Jenga enables practical, high-quality video generation on modern hardware by reducing inference time from minutes to seconds -- without requiring model retraining. Code: https://github.com/dvlab-research/Jenga
format Preprint
id arxiv_https___arxiv_org_abs_2505_16864
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Training-Free Efficient Video Generation via Dynamic Token Carving
Zhang, Yuechen
Xing, Jinbo
Xia, Bin
Liu, Shaoteng
Peng, Bohao
Tao, Xin
Wan, Pengfei
Lo, Eric
Jia, Jiaya
Computer Vision and Pattern Recognition
Despite the remarkable generation quality of video Diffusion Transformer (DiT) models, their practical deployment is severely hindered by extensive computational requirements. This inefficiency stems from two key challenges: the quadratic complexity of self-attention with respect to token length and the multi-step nature of diffusion models. To address these limitations, we present Jenga, a novel inference pipeline that combines dynamic attention carving with progressive resolution generation. Our approach leverages two key insights: (1) early denoising steps do not require high-resolution latents, and (2) later steps do not require dense attention. Jenga introduces a block-wise attention mechanism that dynamically selects relevant token interactions using 3D space-filling curves, alongside a progressive resolution strategy that gradually increases latent resolution during generation. Experimental results demonstrate that Jenga achieves substantial speedups across multiple state-of-the-art video diffusion models while maintaining comparable generation quality (8.83$\times$ speedup with 0.01\% performance drop on VBench). As a plug-and-play solution, Jenga enables practical, high-quality video generation on modern hardware by reducing inference time from minutes to seconds -- without requiring model retraining. Code: https://github.com/dvlab-research/Jenga
title Training-Free Efficient Video Generation via Dynamic Token Carving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.16864