Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yu, Jiwen, Bai, Jianhong, Qin, Yiran, Liu, Quande, Wang, Xintao, Wan, Pengfei, Zhang, Di, Liu, Xihui
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909733174116352
author Yu, Jiwen
Bai, Jianhong
Qin, Yiran
Liu, Quande
Wang, Xintao
Wan, Pengfei
Zhang, Di
Liu, Xihui
author_facet Yu, Jiwen
Bai, Jianhong
Qin, Yiran
Liu, Quande
Wang, Xintao
Wan, Pengfei
Zhang, Di
Liu, Xihui
contents Recent advances in interactive video generation have shown promising results, yet existing approaches struggle with scene-consistent memory capabilities in long video generation due to limited use of historical context. In this work, we propose Context-as-Memory, which utilizes historical context as memory for video generation. It includes two simple yet effective designs: (1) storing context in frame format without additional post-processing; (2) conditioning by concatenating context and frames to be predicted along the frame dimension at the input, requiring no external control modules. Furthermore, considering the enormous computational overhead of incorporating all historical context, we propose the Memory Retrieval module to select truly relevant context frames by determining FOV (Field of View) overlap between camera poses, which significantly reduces the number of candidate frames without substantial information loss. Experiments demonstrate that Context-as-Memory achieves superior memory capabilities in interactive long video generation compared to SOTAs, even generalizing effectively to open-domain scenarios not seen during training. The link of our project page is https://context-as-memory.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2506_03141
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
Yu, Jiwen
Bai, Jianhong
Qin, Yiran
Liu, Quande
Wang, Xintao
Wan, Pengfei
Zhang, Di
Liu, Xihui
Computer Vision and Pattern Recognition
Recent advances in interactive video generation have shown promising results, yet existing approaches struggle with scene-consistent memory capabilities in long video generation due to limited use of historical context. In this work, we propose Context-as-Memory, which utilizes historical context as memory for video generation. It includes two simple yet effective designs: (1) storing context in frame format without additional post-processing; (2) conditioning by concatenating context and frames to be predicted along the frame dimension at the input, requiring no external control modules. Furthermore, considering the enormous computational overhead of incorporating all historical context, we propose the Memory Retrieval module to select truly relevant context frames by determining FOV (Field of View) overlap between camera poses, which significantly reduces the number of candidate frames without substantial information loss. Experiments demonstrate that Context-as-Memory achieves superior memory capabilities in interactive long video generation compared to SOTAs, even generalizing effectively to open-domain scenarios not seen during training. The link of our project page is https://context-as-memory.github.io/.
title Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.03141