REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Thompson, Jacob, Garcia-Lopez, Emiliano, Bisk, Yonatan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909934683160576
author Thompson, Jacob
Garcia-Lopez, Emiliano
Bisk, Yonatan
author_facet Thompson, Jacob
Garcia-Lopez, Emiliano
Bisk, Yonatan
contents Humans build viewpoint-independent cognitive maps through navigation, enabling intuitive reasoning about object permanence and spatial relations. We argue that multimodal large language models (MLLMs), despite extensive video training, lack this fundamental spatial reasoning capability, a critical limitation for embodied applications. To demonstrate these limitations and drive research, we introduce REM (Reasoning over Embodied Multi-Frame Trajectories), a benchmark using controllable 3D environments for long-horizon embodied spatial reasoning. REM systematically evaluates key aspects like object permanence/distinction, spatial relationships, and numerical tracking across dynamic embodied viewpoints. Our evaluation shows that the best-performing current models exhibit promising overall performance, but become increasingly unreliable at even moderate complexity levels easily handled by humans. These findings highlight challenges MLLMs face in developing robust spatial representations from sequential visual input. Consequently, REM provides targeted metrics and diagnostics to foster improved spatial understanding in future models.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00736
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories
Thompson, Jacob
Garcia-Lopez, Emiliano
Bisk, Yonatan
Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
Humans build viewpoint-independent cognitive maps through navigation, enabling intuitive reasoning about object permanence and spatial relations. We argue that multimodal large language models (MLLMs), despite extensive video training, lack this fundamental spatial reasoning capability, a critical limitation for embodied applications. To demonstrate these limitations and drive research, we introduce REM (Reasoning over Embodied Multi-Frame Trajectories), a benchmark using controllable 3D environments for long-horizon embodied spatial reasoning. REM systematically evaluates key aspects like object permanence/distinction, spatial relationships, and numerical tracking across dynamic embodied viewpoints. Our evaluation shows that the best-performing current models exhibit promising overall performance, but become increasingly unreliable at even moderate complexity levels easily handled by humans. These findings highlight challenges MLLMs face in developing robust spatial representations from sequential visual input. Consequently, REM provides targeted metrics and diagnostics to foster improved spatial understanding in future models.
title REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories
topic Machine Learning
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.00736