MIND: Benchmarking Memory Consistency and Action Control in World Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ye, Yixuan, Lu, Xuanyu, Jiang, Yuxin, Gu, Yuchao, Zhao, Rui, Liang, Qiwei, Pan, Jiachun, Zhang, Fengda, Wu, Weijia, Wang, Alex Jinpeng
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917268005322752
author Ye, Yixuan
Lu, Xuanyu
Jiang, Yuxin
Gu, Yuchao
Zhao, Rui
Liang, Qiwei
Pan, Jiachun
Zhang, Fengda
Wu, Weijia
Wang, Alex Jinpeng
author_facet Ye, Yixuan
Lu, Xuanyu
Jiang, Yuxin
Gu, Yuchao
Zhao, Rui
Liang, Qiwei
Pan, Jiachun
Zhang, Fengda
Wu, Weijia
Wang, Alex Jinpeng
contents World models aim to understand, remember, and predict dynamic visual environments, yet a unified benchmark for evaluating their fundamental abilities remains lacking. To address this gap, we introduce MIND, the first open-domain closed-loop revisited benchmark for evaluating Memory consIstency and action coNtrol in worlD models. MIND contains 250 high-quality videos at 1080p and 24 FPS, including 100 (first-person) + 100 (third-person) video clips under a shared action space and 25 + 25 clips across varied action spaces covering eight diverse scenes. We design an efficient evaluation framework to measure two core abilities: memory consistency and action control, capturing temporal stability and contextual coherence across viewpoints. Furthermore, we design various action spaces, including different character movement speeds and camera rotation angles, to evaluate the action generalization capability across different action spaces under shared scenes. To facilitate future performance benchmarking on MIND, we introduce MIND-World, a novel interactive Video-to-World baseline. Extensive experiments demonstrate the completeness of MIND and reveal key challenges in current world models, including the difficulty of maintaining long-term memory consistency and generalizing across action spaces. Code: https://github.com/CSU-JPG/MIND.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08025
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MIND: Benchmarking Memory Consistency and Action Control in World Models
Ye, Yixuan
Lu, Xuanyu
Jiang, Yuxin
Gu, Yuchao
Zhao, Rui
Liang, Qiwei
Pan, Jiachun
Zhang, Fengda
Wu, Weijia
Wang, Alex Jinpeng
Computer Vision and Pattern Recognition
Artificial Intelligence
World models aim to understand, remember, and predict dynamic visual environments, yet a unified benchmark for evaluating their fundamental abilities remains lacking. To address this gap, we introduce MIND, the first open-domain closed-loop revisited benchmark for evaluating Memory consIstency and action coNtrol in worlD models. MIND contains 250 high-quality videos at 1080p and 24 FPS, including 100 (first-person) + 100 (third-person) video clips under a shared action space and 25 + 25 clips across varied action spaces covering eight diverse scenes. We design an efficient evaluation framework to measure two core abilities: memory consistency and action control, capturing temporal stability and contextual coherence across viewpoints. Furthermore, we design various action spaces, including different character movement speeds and camera rotation angles, to evaluate the action generalization capability across different action spaces under shared scenes. To facilitate future performance benchmarking on MIND, we introduce MIND-World, a novel interactive Video-to-World baseline. Extensive experiments demonstrate the completeness of MIND and reveal key challenges in current world models, including the difficulty of maintaining long-term memory consistency and generalizing across action spaces. Code: https://github.com/CSU-JPG/MIND.
title MIND: Benchmarking Memory Consistency and Action Control in World Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.08025