REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories
Fuente:
arXiv
Saved in:
| Main Authors: | Thompson, Jacob, Garcia-Lopez, Emiliano, Bisk, Yonatan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
RAVEN: Resilient Aerial Navigation via Open-Set Semantic Memory and Behavior Adaptation
by: Kim, Seungchan, et al.
Published: (2025)
by: Kim, Seungchan, et al.
Published: (2025)
AmaraSpatial-10K: A Spatially and Semantically Aligned 3D Dataset for Spatial Computing and Embodied AI
by: Salehi, Mohammad Sadegh, et al.
Published: (2026)
by: Salehi, Mohammad Sadegh, et al.
Published: (2026)
VISREAS: Complex Visual Reasoning with Unanswerable Questions
by: Akter, Syeda Nahida, et al.
Published: (2024)
by: Akter, Syeda Nahida, et al.
Published: (2024)
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
by: Li, Chenghao, et al.
Published: (2026)
by: Li, Chenghao, et al.
Published: (2026)
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
by: Yosef, Ron, et al.
Published: (2025)
by: Yosef, Ron, et al.
Published: (2025)
ANAVI: Audio Noise Awareness using Visuals of Indoor environments for NAVIgation
by: Jain, Vidhi, et al.
Published: (2024)
by: Jain, Vidhi, et al.
Published: (2024)
DarkVesselNet: Multi-Modal Remote Sensing and Trajectory Reasoning for Dark Vessel Detection
by: Sharma, Arun
Published: (2026)
by: Sharma, Arun
Published: (2026)
Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning
by: Ganai, Milan, et al.
Published: (2026)
by: Ganai, Milan, et al.
Published: (2026)
Self-Supervised Multi-Frame Neural Scene Flow
by: Liu, Dongrui, et al.
Published: (2024)
by: Liu, Dongrui, et al.
Published: (2024)
3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
by: Ogunleye, Makanjuola, et al.
Published: (2026)
by: Ogunleye, Makanjuola, et al.
Published: (2026)
SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
by: Zhu, Haoyi, et al.
Published: (2024)
by: Zhu, Haoyi, et al.
Published: (2024)
EmboTeam: Grounding LLM Reasoning into Reactive Behavior Trees via PDDL for Embodied Multi-Robot Collaboration
by: Zeng, Haishan, et al.
Published: (2026)
by: Zeng, Haishan, et al.
Published: (2026)
Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks
by: Sharma, Arun
Published: (2026)
by: Sharma, Arun
Published: (2026)
RobotArena $\infty$: Scalable Robot Benchmarking via Real-to-Sim Translation
by: Jangir, Yash, et al.
Published: (2025)
by: Jangir, Yash, et al.
Published: (2025)
Spatial-ViLT: Enhancing Visual Spatial Reasoning through Multi-Task Learning
by: Islam, Chashi Mahiul, et al.
Published: (2025)
by: Islam, Chashi Mahiul, et al.
Published: (2025)
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
by: Wu, Yuqi, et al.
Published: (2024)
by: Wu, Yuqi, et al.
Published: (2024)
From Sparse Decisions to Dense Reasoning: A Multi-attribute Trajectory Paradigm for Multimodal Moderation
by: Gu, Tianle, et al.
Published: (2026)
by: Gu, Tianle, et al.
Published: (2026)
GRADEO: Towards Human-Like Evaluation for Text-to-Video Generation via Multi-Step Reasoning
by: Mou, Zhun, et al.
Published: (2025)
by: Mou, Zhun, et al.
Published: (2025)
LogTinyLLM: Tiny Large Language Models Based Contextual Log Anomaly Detection
by: Ocansey, Isaiah Thompson, et al.
Published: (2025)
by: Ocansey, Isaiah Thompson, et al.
Published: (2025)
MotIF: Motion Instruction Fine-tuning
by: Hwang, Minyoung, et al.
Published: (2024)
by: Hwang, Minyoung, et al.
Published: (2024)
Lifting Embodied World Models for Planning and Control
by: Wang, Alex N., et al.
Published: (2026)
by: Wang, Alex N., et al.
Published: (2026)
Conformal Trajectory Prediction with Multi-View Data Integration in Cooperative Driving
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
VideoPhy: Evaluating Physical Commonsense for Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Beyond a Single Frame: Multi-Frame Spatially Grounded Reasoning Across Volumetric MRI
by: Moukheiber, Lama, et al.
Published: (2026)
by: Moukheiber, Lama, et al.
Published: (2026)
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
by: Li, Chengzu, et al.
Published: (2026)
by: Li, Chengzu, et al.
Published: (2026)
Den-TP: A Density-Balanced Data Curation and Evaluation Framework for Trajectory Prediction
by: Yang, Ruining, et al.
Published: (2024)
by: Yang, Ruining, et al.
Published: (2024)
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
by: Yuan, Yuqian, et al.
Published: (2024)
by: Yuan, Yuqian, et al.
Published: (2024)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
ESCA: Contextualizing Embodied Agents via Scene-Graph Generation
by: Huang, Jiani, et al.
Published: (2025)
by: Huang, Jiani, et al.
Published: (2025)
Learning Visual Abstract Reasoning through Dual-Stream Networks
by: Zhao, Kai, et al.
Published: (2024)
by: Zhao, Kai, et al.
Published: (2024)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
by: Batra, Hunar, et al.
Published: (2025)
by: Batra, Hunar, et al.
Published: (2025)
CXR-TFT: Multi-Modal Temporal Fusion Transformer for Predicting Chest X-ray Trajectories
by: Arora, Mehak, et al.
Published: (2025)
by: Arora, Mehak, et al.
Published: (2025)
Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion
by: Kim, Dongjun, et al.
Published: (2023)
by: Kim, Dongjun, et al.
Published: (2023)
Multi-Label Contrastive Learning for Abstract Visual Reasoning
by: Małkiński, Mikołaj, et al.
Published: (2020)
by: Małkiński, Mikołaj, et al.
Published: (2020)
Spatially-Aware Evaluation of Segmentation Uncertainty
by: Zeevi, Tal, et al.
Published: (2025)
by: Zeevi, Tal, et al.
Published: (2025)
Knot So Simple: A Minimalistic Environment for Spatial Reasoning
by: Chen, Zizhao, et al.
Published: (2025)
by: Chen, Zizhao, et al.
Published: (2025)
Similar Items
-
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
by: Hu, Wenbo, et al.
Published: (2025) -
RAVEN: Resilient Aerial Navigation via Open-Set Semantic Memory and Behavior Adaptation
by: Kim, Seungchan, et al.
Published: (2025) -
AmaraSpatial-10K: A Spatially and Semantically Aligned 3D Dataset for Spatial Computing and Embodied AI
by: Salehi, Mohammad Sadegh, et al.
Published: (2026) -
VISREAS: Complex Visual Reasoning with Unanswerable Questions
by: Akter, Syeda Nahida, et al.
Published: (2024) -
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
by: NVIDIA, et al.
Published: (2025)