Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Kaijin, Liang, Dingkang, Zhou, Xin, Ding, Yikang, Liu, Xiaoqiang, Wan, Pengfei, Bai, Xiang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915896408145920
author Chen, Kaijin
Liang, Dingkang
Zhou, Xin
Ding, Yikang
Liu, Xiaoqiang
Wan, Pengfei
Bai, Xiang
author_facet Chen, Kaijin
Liang, Dingkang
Zhou, Xin
Ding, Yikang
Liu, Xiaoqiang
Wan, Pengfei
Bai, Xiang
contents Video world models have shown immense potential in simulating the physical world, yet existing memory mechanisms primarily treat environments as static canvases. When dynamic subjects hide out of sight and later re-emerge, current methods often struggle, leading to frozen, distorted, or vanishing subjects. To address this, we introduce Hybrid Memory, a novel paradigm requiring models to simultaneously act as precise archivists for static backgrounds and vigilant trackers for dynamic subjects, ensuring motion continuity during out-of-view intervals. To facilitate research in this direction, we construct HM-World, the first large-scale video dataset dedicated to hybrid memory. It features 59K high-fidelity clips with decoupled camera and subject trajectories, encompassing 17 diverse scenes, 49 distinct subjects, and meticulously designed exit-entry events to rigorously evaluate hybrid coherence. Furthermore, we propose HyDRA, a specialized memory architecture that compresses memory into tokens and utilizes a spatiotemporal relevance-driven retrieval mechanism. By selectively attending to relevant motion cues, HyDRA effectively preserves the identity and motion of hidden subjects. Extensive experiments on HM-World demonstrate that our method significantly outperforms state-of-the-art approaches in both dynamic subject consistency and overall generation quality. Code is publicly available at https://github.com/H-EmbodVis/HyDRA.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25716
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
Chen, Kaijin
Liang, Dingkang
Zhou, Xin
Ding, Yikang
Liu, Xiaoqiang
Wan, Pengfei
Bai, Xiang
Computer Vision and Pattern Recognition
Artificial Intelligence
Video world models have shown immense potential in simulating the physical world, yet existing memory mechanisms primarily treat environments as static canvases. When dynamic subjects hide out of sight and later re-emerge, current methods often struggle, leading to frozen, distorted, or vanishing subjects. To address this, we introduce Hybrid Memory, a novel paradigm requiring models to simultaneously act as precise archivists for static backgrounds and vigilant trackers for dynamic subjects, ensuring motion continuity during out-of-view intervals. To facilitate research in this direction, we construct HM-World, the first large-scale video dataset dedicated to hybrid memory. It features 59K high-fidelity clips with decoupled camera and subject trajectories, encompassing 17 diverse scenes, 49 distinct subjects, and meticulously designed exit-entry events to rigorously evaluate hybrid coherence. Furthermore, we propose HyDRA, a specialized memory architecture that compresses memory into tokens and utilizes a spatiotemporal relevance-driven retrieval mechanism. By selectively attending to relevant motion cues, HyDRA effectively preserves the identity and motion of hidden subjects. Extensive experiments on HM-World demonstrate that our method significantly outperforms state-of-the-art approaches in both dynamic subject consistency and overall generation quality. Code is publicly available at https://github.com/H-EmbodVis/HyDRA.
title Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2603.25716