Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Luo, Jiayi, Liu, Qiyan, Wang, Tengyang, Liu, JunHao, Chen, Jiayu, Wang, Cong, Zhu, Hanxin, Gao, Chen, Hu, Xiaobin, Sun, Qingyun, Chen, Zhibo
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917544320827392
author Luo, Jiayi
Liu, Qiyan
Wang, Tengyang
Liu, JunHao
Chen, Jiayu
Wang, Cong
Zhu, Hanxin
Gao, Chen
Hu, Xiaobin
Sun, Qingyun
Chen, Zhibo
author_facet Luo, Jiayi
Liu, Qiyan
Wang, Tengyang
Liu, JunHao
Chen, Jiayu
Wang, Cong
Zhu, Hanxin
Gao, Chen
Hu, Xiaobin
Sun, Qingyun
Chen, Zhibo
contents Autoregressive (AR) video generation has emerged as a promising paradigm for long-horizon video synthesis, where each frame is generated conditioned on previously generated tokens. To accelerate inference, the KV cache is used to avoid redundant recomputation across generation steps. Nevertheless, its growth with generation length introduces increasing memory and error accumulation, limiting the scalability of AR models to even longer sequences. Existing KV cache compression methods mitigate this issue by selectively retaining only video tokens deemed important. However, most existing methods assess token importance using short-horizon signals derived from the current or historical generation context, making these methods prone to overlooking tokens that appear unimportant at early steps but later become critical for future frames. In this work, we identify an important property of trained AR video models: although RoPE-modulated queries evolve across autoregressive steps, the underlying canonical pre-RoPE query distribution remains remarkably stable throughout the video generation process. This approximate stationarity implies that future query distributions are estimable from historical statistics, enabling principled future-aware cache decisions without any additional training. Building on this insight, we propose Future Forcing, a training-free future-aware KV cache policy for AR video generation. Specifically, Future Forcing first constructs a future query proxy from historical statistics, then scores KV cache tokens by their importance under this proxy, and finally merges redundant token pairs within the affine subspace induced by the future query. Extensive experiments show that Future Forcing improves long-horizon consistency under limited KV caches, achieving up to 1.49 improvement in subject consistency on VBench-Long for 60s generation over existing AR video KV cache policies.
format Preprint
id arxiv_https___arxiv_org_abs_2605_30083
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
Luo, Jiayi
Liu, Qiyan
Wang, Tengyang
Liu, JunHao
Chen, Jiayu
Wang, Cong
Zhu, Hanxin
Gao, Chen
Hu, Xiaobin
Sun, Qingyun
Chen, Zhibo
Computer Vision and Pattern Recognition
Autoregressive (AR) video generation has emerged as a promising paradigm for long-horizon video synthesis, where each frame is generated conditioned on previously generated tokens. To accelerate inference, the KV cache is used to avoid redundant recomputation across generation steps. Nevertheless, its growth with generation length introduces increasing memory and error accumulation, limiting the scalability of AR models to even longer sequences. Existing KV cache compression methods mitigate this issue by selectively retaining only video tokens deemed important. However, most existing methods assess token importance using short-horizon signals derived from the current or historical generation context, making these methods prone to overlooking tokens that appear unimportant at early steps but later become critical for future frames. In this work, we identify an important property of trained AR video models: although RoPE-modulated queries evolve across autoregressive steps, the underlying canonical pre-RoPE query distribution remains remarkably stable throughout the video generation process. This approximate stationarity implies that future query distributions are estimable from historical statistics, enabling principled future-aware cache decisions without any additional training. Building on this insight, we propose Future Forcing, a training-free future-aware KV cache policy for AR video generation. Specifically, Future Forcing first constructs a future query proxy from historical statistics, then scores KV cache tokens by their importance under this proxy, and finally merges redundant token pairs within the affine subspace induced by the future query. Extensive experiments show that Future Forcing improves long-horizon consistency under limited KV caches, achieving up to 1.49 improvement in subject consistency on VBench-Long for 60s generation over existing AR video KV cache policies.
title Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.30083