Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yi, Jung, Jang, Wooseok, Cho, Paul Hyunbin, Nam, Jisu, Yoon, Heeji, Kim, Seungryong
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908694357213184
author Yi, Jung
Jang, Wooseok
Cho, Paul Hyunbin
Nam, Jisu
Yoon, Heeji
Kim, Seungryong
author_facet Yi, Jung
Jang, Wooseok
Cho, Paul Hyunbin
Nam, Jisu
Yoon, Heeji
Kim, Seungryong
contents Recent advances in autoregressive video diffusion have enabled real-time frame streaming, yet existing solutions still suffer from temporal repetition, drift, and motion deceleration. We find that naively applying StreamingLLM-style attention sinks to video diffusion leads to fidelity degradation and motion stagnation. To overcome this, we introduce Deep Forcing, which consists of two training-free mechanisms that address this without any fine-tuning. Specifically, 1) Deep Sink dedicates half of the sliding window to persistent sink tokens and re-aligns their temporal RoPE phase to the current timeline, stabilizing global context during long rollouts. 2) Participative Compression performs importance-aware KV cache pruning that preserves only tokens actively participating in recent attention while safely discarding redundant and degraded history, minimizing error accumulation under out-of-distribution length generation. Together, these components enable over 12x extrapolation (e.g. 5s-trained to 60s+ generation) with better imaging quality than LongLive, better aesthetic quality than RollingForcing, almost maintaining overall consistency, and substantial gains in dynamic degree, all while maintaining real-time generation. Our results demonstrate that training-free KV-cache management can match or exceed training-based approaches for autoregressively streaming long-video generation.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05081
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression
Yi, Jung
Jang, Wooseok
Cho, Paul Hyunbin
Nam, Jisu
Yoon, Heeji
Kim, Seungryong
Computer Vision and Pattern Recognition
Recent advances in autoregressive video diffusion have enabled real-time frame streaming, yet existing solutions still suffer from temporal repetition, drift, and motion deceleration. We find that naively applying StreamingLLM-style attention sinks to video diffusion leads to fidelity degradation and motion stagnation. To overcome this, we introduce Deep Forcing, which consists of two training-free mechanisms that address this without any fine-tuning. Specifically, 1) Deep Sink dedicates half of the sliding window to persistent sink tokens and re-aligns their temporal RoPE phase to the current timeline, stabilizing global context during long rollouts. 2) Participative Compression performs importance-aware KV cache pruning that preserves only tokens actively participating in recent attention while safely discarding redundant and degraded history, minimizing error accumulation under out-of-distribution length generation. Together, these components enable over 12x extrapolation (e.g. 5s-trained to 60s+ generation) with better imaging quality than LongLive, better aesthetic quality than RollingForcing, almost maintaining overall consistency, and substantial gains in dynamic degree, all while maintaining real-time generation. Our results demonstrate that training-free KV-cache management can match or exceed training-based approaches for autoregressively streaming long-video generation.
title Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.05081