Saved in:
Bibliographic Details
Main Authors: Lv, Chengtao, Shi, Yumeng, Huang, Yushi, Gong, Ruihao, Ren, Shen, Wang, Wenya
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.04789
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911712362364928
author Lv, Chengtao
Shi, Yumeng
Huang, Yushi
Gong, Ruihao
Ren, Shen
Wang, Wenya
author_facet Lv, Chengtao
Shi, Yumeng
Huang, Yushi
Gong, Ruihao
Ren, Shen
Wang, Wenya
contents Advanced autoregressive (AR) video generation models have improved visual fidelity and interactivity, but the quadratic complexity of attention remains a primary bottleneck for efficient deployment. While existing sparse attention solutions have shown promise on bidirectional models, we identify that applying these solutions to AR models leads to considerable performance degradation for two reasons: isolated consideration of chunk generation and insufficient utilization of past informative context. Motivated by these observations, we propose \textsc{Light Forcing}, the \textit{first} sparse attention solution tailored for AR video generation models. It incorporates a \textit{Chunk-Aware Growth} mechanism to quantitatively estimate the contribution of each chunk, which determines their sparsity allocation. This progressive sparsity increase strategy enables the current chunk to inherit prior knowledge in earlier chunks during generation. Additionally, we introduce a \textit{Hierarchical Sparse Attention} to capture informative historical and local context in a coarse-to-fine manner. Such two-level mask selection strategy (\ie, frame and block level) can adaptively handle diverse attention patterns. Extensive experiments demonstrate that our method outperforms existing sparse attention in quality (\eg, 84.5 on VBench) and efficiency (\eg, $1.2{\sim}1.3\times$ end-to-end speedup). Combined with FP8 quantization and LightVAE, \textsc{Light Forcing} further achieves a $2.3\times$ speedup and 19.7\,FPS on an RTX~5090 GPU. Code will be released at \href{https://github.com/chengtao-lv/LightForcing}{https://github.com/chengtao-lv/LightForcing}.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04789
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention
Lv, Chengtao
Shi, Yumeng
Huang, Yushi
Gong, Ruihao
Ren, Shen
Wang, Wenya
Computer Vision and Pattern Recognition
Advanced autoregressive (AR) video generation models have improved visual fidelity and interactivity, but the quadratic complexity of attention remains a primary bottleneck for efficient deployment. While existing sparse attention solutions have shown promise on bidirectional models, we identify that applying these solutions to AR models leads to considerable performance degradation for two reasons: isolated consideration of chunk generation and insufficient utilization of past informative context. Motivated by these observations, we propose \textsc{Light Forcing}, the \textit{first} sparse attention solution tailored for AR video generation models. It incorporates a \textit{Chunk-Aware Growth} mechanism to quantitatively estimate the contribution of each chunk, which determines their sparsity allocation. This progressive sparsity increase strategy enables the current chunk to inherit prior knowledge in earlier chunks during generation. Additionally, we introduce a \textit{Hierarchical Sparse Attention} to capture informative historical and local context in a coarse-to-fine manner. Such two-level mask selection strategy (\ie, frame and block level) can adaptively handle diverse attention patterns. Extensive experiments demonstrate that our method outperforms existing sparse attention in quality (\eg, 84.5 on VBench) and efficiency (\eg, $1.2{\sim}1.3\times$ end-to-end speedup). Combined with FP8 quantization and LightVAE, \textsc{Light Forcing} further achieves a $2.3\times$ speedup and 19.7\,FPS on an RTX~5090 GPU. Code will be released at \href{https://github.com/chengtao-lv/LightForcing}{https://github.com/chengtao-lv/LightForcing}.
title Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.04789