LoL: Longer than Longer, Scaling Video Generation to Hour

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Cui, Justin, Wu, Jie, Li, Ming, Yang, Tao, Li, Xiaojie, Wang, Rui, Bai, Andrew, Ban, Yuanhao, Hsieh, Cho-Jui
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917219661774848
author Cui, Justin
Wu, Jie
Li, Ming
Yang, Tao
Li, Xiaojie
Wang, Rui
Bai, Andrew
Ban, Yuanhao
Hsieh, Cho-Jui
author_facet Cui, Justin
Wu, Jie
Li, Ming
Yang, Tao
Li, Xiaojie
Wang, Rui
Bai, Andrew
Ban, Yuanhao
Hsieh, Cho-Jui
contents Recent research in long-form video generation has shifted from bidirectional to autoregressive models, yet these methods commonly suffer from error accumulation and a loss of long-term coherence. While attention sink frames have been introduced to mitigate this performance decay, they often induce a critical failure mode we term sink-collapse: the generated content repeatedly reverts to the sink frame, resulting in abrupt scene resets and cyclic motion patterns. Our analysis reveals that sink-collapse originates from an inherent conflict between the periodic structure of Rotary Position Embedding (RoPE) and the multi-head attention mechanisms prevalent in current generative models. To address it, we propose a lightweight, training-free approach that effectively suppresses this behavior by introducing multi-head RoPE jitter that breaks inter-head attention homogenization and mitigates long-horizon collapse. Extensive experiments show that our method successfully alleviates sink-collapse while preserving generation quality. To the best of our knowledge, this work achieves the first demonstration of real-time, streaming, and infinite-length video generation with little quality decay. As an illustration of this robustness, we generate continuous videos up to 12 hours in length, which, to our knowledge, is among the longest publicly demonstrated results in streaming video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2601_16914
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LoL: Longer than Longer, Scaling Video Generation to Hour
Cui, Justin
Wu, Jie
Li, Ming
Yang, Tao
Li, Xiaojie
Wang, Rui
Bai, Andrew
Ban, Yuanhao
Hsieh, Cho-Jui
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent research in long-form video generation has shifted from bidirectional to autoregressive models, yet these methods commonly suffer from error accumulation and a loss of long-term coherence. While attention sink frames have been introduced to mitigate this performance decay, they often induce a critical failure mode we term sink-collapse: the generated content repeatedly reverts to the sink frame, resulting in abrupt scene resets and cyclic motion patterns. Our analysis reveals that sink-collapse originates from an inherent conflict between the periodic structure of Rotary Position Embedding (RoPE) and the multi-head attention mechanisms prevalent in current generative models. To address it, we propose a lightweight, training-free approach that effectively suppresses this behavior by introducing multi-head RoPE jitter that breaks inter-head attention homogenization and mitigates long-horizon collapse. Extensive experiments show that our method successfully alleviates sink-collapse while preserving generation quality. To the best of our knowledge, this work achieves the first demonstration of real-time, streaming, and infinite-length video generation with little quality decay. As an illustration of this robustness, we generate continuous videos up to 12 hours in length, which, to our knowledge, is among the longest publicly demonstrated results in streaming video generation.
title LoL: Longer than Longer, Scaling Video Generation to Hour
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2601.16914