Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Kaijin, Liang, Dingkang, Zhou, Xin, Ding, Yikang, Liu, Xiaoqiang, Wan, Pengfei, Bai, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models
by: Ma, Ziqi, et al.
Published: (2026)
by: Ma, Ziqi, et al.
Published: (2026)
LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
by: Duan, Zicheng, et al.
Published: (2026)
by: Duan, Zicheng, et al.
Published: (2026)
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
by: Plizzari, Chiara, et al.
Published: (2024)
by: Plizzari, Chiara, et al.
Published: (2024)
Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution
by: Xu, Tianshuo, et al.
Published: (2026)
by: Xu, Tianshuo, et al.
Published: (2026)
HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Out of Sight, Still in Mind: Reasoning and Planning about Unobserved Objects with Video Tracking Enabled Memory Models
by: Huang, Yixuan, et al.
Published: (2023)
by: Huang, Yixuan, et al.
Published: (2023)
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
by: Sun, Zhengyang, et al.
Published: (2026)
by: Sun, Zhengyang, et al.
Published: (2026)
HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation
by: Zhou, Xin, et al.
Published: (2026)
by: Zhou, Xin, et al.
Published: (2026)
The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey
by: Tu, Sifan, et al.
Published: (2025)
by: Tu, Sifan, et al.
Published: (2025)
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
by: Ravi, Sahithya, et al.
Published: (2025)
by: Ravi, Sahithya, et al.
Published: (2025)
More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models
by: Lin, Hongkai, et al.
Published: (2025)
by: Lin, Hongkai, et al.
Published: (2025)
PointTPA: Dynamic Network Parameter Adaptation for 3D Scene Understanding
by: Liu, Siyuan, et al.
Published: (2026)
by: Liu, Siyuan, et al.
Published: (2026)
Geometry-Aware Implicit Memory for Video World Models
by: Wei, Zhengxuan, et al.
Published: (2026)
by: Wei, Zhengxuan, et al.
Published: (2026)
Out of Sight, Out of Track: Adversarial Attacks on Propagation-based Multi-Object Trackers via Query State Manipulation
by: Bouzidi, Halima, et al.
Published: (2026)
by: Bouzidi, Halima, et al.
Published: (2026)
A Unified Image-Dense Annotation Generation Model for Underwater Scenes
by: Lin, Hongkai, et al.
Published: (2025)
by: Lin, Hongkai, et al.
Published: (2025)
Dynamic Adapter Meets Prompt Tuning: Parameter-Efficient Transfer Learning for Point Cloud Analysis
by: Zhou, Xin, et al.
Published: (2024)
by: Zhou, Xin, et al.
Published: (2024)
MosaicMem: Hybrid Spatial Memory for Controllable Video World Models
by: Yu, Wei, et al.
Published: (2026)
by: Yu, Wei, et al.
Published: (2026)
UniFuture: A 4D Driving World Model for Future Generation and Perception
by: Liang, Dingkang, et al.
Published: (2025)
by: Liang, Dingkang, et al.
Published: (2025)
Parameter-Efficient Fine-Tuning in Spectral Domain for Point Cloud Learning
by: Liang, Dingkang, et al.
Published: (2024)
by: Liang, Dingkang, et al.
Published: (2024)
A Unified Framework for 3D Scene Understanding
by: Xu, Wei, et al.
Published: (2024)
by: Xu, Wei, et al.
Published: (2024)
MINIMA: Modality Invariant Image Matching
by: Ren, Jiangwei, et al.
Published: (2024)
by: Ren, Jiangwei, et al.
Published: (2024)
ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos
by: Wang, Xiaodong, et al.
Published: (2025)
by: Wang, Xiaodong, et al.
Published: (2025)
Segment Every Out-of-Distribution Object
by: Zhao, Wenjie, et al.
Published: (2023)
by: Zhao, Wenjie, et al.
Published: (2023)
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
Anomaly Detection by Adapting a pre-trained Vision Language Model
by: Cai, Yuxuan, et al.
Published: (2024)
by: Cai, Yuxuan, et al.
Published: (2024)
MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
by: Ji, Sihui, et al.
Published: (2025)
by: Ji, Sihui, et al.
Published: (2025)
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
by: Zhao, Zongchuang, et al.
Published: (2025)
by: Zhao, Zongchuang, et al.
Published: (2025)
PointMamba: A Simple State Space Model for Point Cloud Analysis
by: Liang, Dingkang, et al.
Published: (2024)
by: Liang, Dingkang, et al.
Published: (2024)
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
DynProto: Dynamic Prototype Evolution for Out-of-Distribution Detection
by: Wu, Yanqi, et al.
Published: (2026)
by: Wu, Yanqi, et al.
Published: (2026)
Out-of-Distribution Detection using Neural Activation Prior
by: Wan, Weilin, et al.
Published: (2024)
by: Wan, Weilin, et al.
Published: (2024)
Towards Generalizable Robotic Manipulation in Dynamic Environments
by: Fang, Heng, et al.
Published: (2026)
by: Fang, Heng, et al.
Published: (2026)
FakeOut: Leveraging Out-of-domain Self-supervision for Multi-modal Video Deepfake Detection
by: Knafo, Gil, et al.
Published: (2022)
by: Knafo, Gil, et al.
Published: (2022)
OOSTraj: Out-of-Sight Trajectory Prediction With Vision-Positioning Denoising
by: Zhang, Haichao, et al.
Published: (2024)
by: Zhang, Haichao, et al.
Published: (2024)
Frame In-N-Out: Unbounded Controllable Image-to-Video Generation
by: Wang, Boyang, et al.
Published: (2025)
by: Wang, Boyang, et al.
Published: (2025)
Efficient Video Diffusion Models: Advancements and Challenges
by: Shao, Shitong, et al.
Published: (2026)
by: Shao, Shitong, et al.
Published: (2026)
Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
by: Huang, Mingxin, et al.
Published: (2024)
by: Huang, Mingxin, et al.
Published: (2024)
A Mechanistic View on Video Generation as World Models: State and Dynamics
by: Wang, Luozhou, et al.
Published: (2026)
by: Wang, Luozhou, et al.
Published: (2026)
Similar Items
-
Out of Sight, Out of Mind? Evaluating State Evolution in Video World Models
by: Ma, Ziqi, et al.
Published: (2026) -
LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models
by: Duan, Zicheng, et al.
Published: (2026) -
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
by: Plizzari, Chiara, et al.
Published: (2024) -
Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching
by: Zhou, Xin, et al.
Published: (2025) -
Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution
by: Xu, Tianshuo, et al.
Published: (2026)