ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911287422746624 |
|---|---|
| author | Wang, Qineng Huang, Wenlong Zhou, Yu Yin, Hang Bao, Tianwei Lyu, Jianwen Liu, Weiyu Zhang, Ruohan Wu, Jiajun Fei-Fei, Li Li, Manling |
| author_facet | Wang, Qineng Huang, Wenlong Zhou, Yu Yin, Hang Bao, Tianwei Lyu, Jianwen Liu, Weiyu Zhang, Ruohan Wu, Jiajun Fei-Fei, Li Li, Manling |
| contents | Embodied cognition argues that intelligence arises from sensorimotor interaction rather than passive observation. It raises an intriguing question: do modern vision-language models (VLMs), trained largely in a disembodied manner, exhibit signs of embodied cognition? We introduce ENACT, a benchmark that casts evaluation of embodied cognition as world modeling from egocentric interaction in a visual question answering (VQA) format. Framed as a partially observable Markov decision process (POMDP) whose actions are scene graph changes, ENACT comprises two complementary sequence reordering tasks: forward world modeling (reorder shuffled observations given actions) and inverse world modeling (reorder shuffled actions given observations). While conceptually simple, solving these tasks implicitly demands capabilities central to embodied cognition-affordance recognition, action-effect reasoning, embodied awareness, and interactive, long-horizon memory from partially observable egocentric input, while avoiding low-level image synthesis that could confound the evaluation. We provide a scalable pipeline that synthesizes QA pairs from robotics simulation (BEHAVIOR) and evaluates models on 8,972 QA pairs spanning long-horizon home-scale activities. Experiments reveal a performance gap between frontier VLMs and humans that widens with interaction horizon. Models consistently perform better on the inverse task than the forward one and exhibit anthropocentric biases, including a preference for right-handed actions and degradation when camera intrinsics or viewpoints deviate from human vision. Website at https://enact-embodied-cognition.github.io/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_20937 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction Wang, Qineng Huang, Wenlong Zhou, Yu Yin, Hang Bao, Tianwei Lyu, Jianwen Liu, Weiyu Zhang, Ruohan Wu, Jiajun Fei-Fei, Li Li, Manling Artificial Intelligence Computation and Language Computer Vision and Pattern Recognition Robotics Embodied cognition argues that intelligence arises from sensorimotor interaction rather than passive observation. It raises an intriguing question: do modern vision-language models (VLMs), trained largely in a disembodied manner, exhibit signs of embodied cognition? We introduce ENACT, a benchmark that casts evaluation of embodied cognition as world modeling from egocentric interaction in a visual question answering (VQA) format. Framed as a partially observable Markov decision process (POMDP) whose actions are scene graph changes, ENACT comprises two complementary sequence reordering tasks: forward world modeling (reorder shuffled observations given actions) and inverse world modeling (reorder shuffled actions given observations). While conceptually simple, solving these tasks implicitly demands capabilities central to embodied cognition-affordance recognition, action-effect reasoning, embodied awareness, and interactive, long-horizon memory from partially observable egocentric input, while avoiding low-level image synthesis that could confound the evaluation. We provide a scalable pipeline that synthesizes QA pairs from robotics simulation (BEHAVIOR) and evaluates models on 8,972 QA pairs spanning long-horizon home-scale activities. Experiments reveal a performance gap between frontier VLMs and humans that widens with interaction horizon. Models consistently perform better on the inverse task than the forward one and exhibit anthropocentric biases, including a preference for right-handed actions and degradation when camera intrinsics or viewpoints deviate from human vision. Website at https://enact-embodied-cognition.github.io/. |
| title | ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction |
| topic | Artificial Intelligence Computation and Language Computer Vision and Pattern Recognition Robotics |
| url | https://arxiv.org/abs/2511.20937 |