MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ji, Sihui, Chen, Xi, Yang, Shuai, Tao, Xin, Wan, Pengfei, Zhao, Hengshuang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909965646561280
author Ji, Sihui
Chen, Xi
Yang, Shuai
Tao, Xin
Wan, Pengfei
Zhao, Hengshuang
author_facet Ji, Sihui
Chen, Xi
Yang, Shuai
Tao, Xin
Wan, Pengfei
Zhao, Hengshuang
contents The core challenge for streaming video generation is maintaining the content consistency in long context, which poses high requirement for the memory design. Most existing solutions maintain the memory by compressing historical frames with predefined strategies. However, different to-generate video chunks should refer to different historical cues, which is hard to satisfy with fixed strategies. In this work, we propose MemFlow to address this problem. Specifically, before generating the coming chunk, we dynamically update the memory bank by retrieving the most relevant historical frames with the text prompt of this chunk. This design enables narrative coherence even if new event happens or scenario switches in future frames. In addition, during generation, we only activate the most relevant tokens in the memory bank for each query in the attention layers, which effectively guarantees the generation efficiency. In this way, MemFlow achieves outstanding long-context consistency with negligible computation burden (7.9% speed reduction compared with the memory-free baseline) and keeps the compatibility with any streaming video generation model with KV cache.
format Preprint
id arxiv_https___arxiv_org_abs_2512_14699
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
Ji, Sihui
Chen, Xi
Yang, Shuai
Tao, Xin
Wan, Pengfei
Zhao, Hengshuang
Computer Vision and Pattern Recognition
The core challenge for streaming video generation is maintaining the content consistency in long context, which poses high requirement for the memory design. Most existing solutions maintain the memory by compressing historical frames with predefined strategies. However, different to-generate video chunks should refer to different historical cues, which is hard to satisfy with fixed strategies. In this work, we propose MemFlow to address this problem. Specifically, before generating the coming chunk, we dynamically update the memory bank by retrieving the most relevant historical frames with the text prompt of this chunk. This design enables narrative coherence even if new event happens or scenario switches in future frames. In addition, during generation, we only activate the most relevant tokens in the memory bank for each query in the attention layers, which effectively guarantees the generation efficiency. In this way, MemFlow achieves outstanding long-context consistency with negligible computation burden (7.9% speed reduction compared with the memory-free baseline) and keeps the compatibility with any streaming video generation model with KV cache.
title MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.14699