Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Tianshuo, Xie, Yichen, Meng, Depu, Peng, Chensheng, Herau, Quentin, Jiang, Bo, Hu, Yihan, Zhan, Wei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916044498534400
author Xu, Tianshuo
Xie, Yichen
Meng, Depu
Peng, Chensheng
Herau, Quentin
Jiang, Bo
Hu, Yihan
Zhan, Wei
author_facet Xu, Tianshuo
Xie, Yichen
Meng, Depu
Peng, Chensheng
Herau, Quentin
Jiang, Bo
Hu, Yihan
Zhan, Wei
contents Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption. This is not simply a capacity problem: pretrained video diffusion transformers already possess KV-cache mechanisms capable of non-local retrieval, but they are rarely trained to use them as dynamic memory. We introduce ReMind, a framework eliciting dynamic memory behavior via memory-oriented data, event-aware training, and cache adaptation. Organized around a taxonomy of 100+ dynamic events, we build a camera-annotated training mixture combining VLM-filtered real videos, generated hard dynamics, synthetic camera loops, and memory-interruption augmentations. Each clip is converted into a frame graph with protected anchors, degraded intervals, and explicit temporal gaps. A node-structured curriculum, including node-drop, noisy memory, frontier continuation, and reference-cache training, forces the model to retrieve relevant past states across interruptions rather than relying solely on local continuity. PM-RoPE, an elegant camera-phase RoPE extension, unlocks spatiotemporal retrieval at a single-attention cost while preserving pretrained pathways. ReMind achieves the best overall scores on STEVO-Bench and recovery tasks. Furthermore, general image-to-video evaluations confirm this curriculum avoids catastrophic forgetting. We will open-source our code, data, and models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25333
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution
Xu, Tianshuo
Xie, Yichen
Meng, Depu
Peng, Chensheng
Herau, Quentin
Jiang, Bo
Hu, Yihan
Zhan, Wei
Computer Vision and Pattern Recognition
Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption. This is not simply a capacity problem: pretrained video diffusion transformers already possess KV-cache mechanisms capable of non-local retrieval, but they are rarely trained to use them as dynamic memory. We introduce ReMind, a framework eliciting dynamic memory behavior via memory-oriented data, event-aware training, and cache adaptation. Organized around a taxonomy of 100+ dynamic events, we build a camera-annotated training mixture combining VLM-filtered real videos, generated hard dynamics, synthetic camera loops, and memory-interruption augmentations. Each clip is converted into a frame graph with protected anchors, degraded intervals, and explicit temporal gaps. A node-structured curriculum, including node-drop, noisy memory, frontier continuation, and reference-cache training, forces the model to retrieve relevant past states across interruptions rather than relying solely on local continuity. PM-RoPE, an elegant camera-phase RoPE extension, unlocks spatiotemporal retrieval at a single-attention cost while preserving pretrained pathways. ReMind achieves the best overall scores on STEVO-Bench and recovery tasks. Furthermore, general image-to-video evaluations confirm this curriculum avoids catastrophic forgetting. We will open-source our code, data, and models.
title Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.25333