How Good are Foundation Models in Step-by-Step Embodied Reasoning?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dissanayake, Dinura, Heakl, Ahmed, Thawakar, Omkar, Ahsan, Noor, Thawkar, Ritesh, More, Ketan, Lahoud, Jean, Anwer, Rao, Cholakkal, Hisham, Laptev, Ivan, Khan, Fahad Shahbaz, Khan, Salman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912598420619264
author Dissanayake, Dinura
Heakl, Ahmed
Thawakar, Omkar
Ahsan, Noor
Thawkar, Ritesh
More, Ketan
Lahoud, Jean
Anwer, Rao
Cholakkal, Hisham
Laptev, Ivan
Khan, Fahad Shahbaz
Khan, Salman
author_facet Dissanayake, Dinura
Heakl, Ahmed
Thawakar, Omkar
Ahsan, Noor
Thawkar, Ritesh
More, Ketan
Lahoud, Jean
Anwer, Rao
Cholakkal, Hisham
Laptev, Ivan
Khan, Fahad Shahbaz
Khan, Salman
contents Embodied agents operating in the physical world must make decisions that are not only effective but also safe, spatially coherent, and grounded in context. While recent advances in large multimodal models (LMMs) have shown promising capabilities in visual understanding and language generation, their ability to perform structured reasoning for real-world embodied tasks remains underexplored. In this work, we aim to understand how well foundation models can perform step-by-step reasoning in embodied environments. To this end, we propose the Foundation Model Embodied Reasoning (FoMER) benchmark, designed to evaluate the reasoning capabilities of LMMs in complex embodied decision-making scenarios. Our benchmark spans a diverse set of tasks that require agents to interpret multimodal observations, reason about physical constraints and safety, and generate valid next actions in natural language. We present (i) a large-scale, curated suite of embodied reasoning tasks, (ii) a novel evaluation framework that disentangles perceptual grounding from action reasoning, and (iii) empirical analysis of several leading LMMs under this setting. Our benchmark includes over 1.1k samples with detailed step-by-step reasoning across 10 tasks and 8 embodiments, covering three different robot types. Our results highlight both the potential and current limitations of LMMs in embodied reasoning, pointing towards key challenges and opportunities for future research in robot intelligence. Our data and code will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15293
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Good are Foundation Models in Step-by-Step Embodied Reasoning?
Dissanayake, Dinura
Heakl, Ahmed
Thawakar, Omkar
Ahsan, Noor
Thawkar, Ritesh
More, Ketan
Lahoud, Jean
Anwer, Rao
Cholakkal, Hisham
Laptev, Ivan
Khan, Fahad Shahbaz
Khan, Salman
Computer Vision and Pattern Recognition
Robotics
Embodied agents operating in the physical world must make decisions that are not only effective but also safe, spatially coherent, and grounded in context. While recent advances in large multimodal models (LMMs) have shown promising capabilities in visual understanding and language generation, their ability to perform structured reasoning for real-world embodied tasks remains underexplored. In this work, we aim to understand how well foundation models can perform step-by-step reasoning in embodied environments. To this end, we propose the Foundation Model Embodied Reasoning (FoMER) benchmark, designed to evaluate the reasoning capabilities of LMMs in complex embodied decision-making scenarios. Our benchmark spans a diverse set of tasks that require agents to interpret multimodal observations, reason about physical constraints and safety, and generate valid next actions in natural language. We present (i) a large-scale, curated suite of embodied reasoning tasks, (ii) a novel evaluation framework that disentangles perceptual grounding from action reasoning, and (iii) empirical analysis of several leading LMMs under this setting. Our benchmark includes over 1.1k samples with detailed step-by-step reasoning across 10 tasks and 8 embodiments, covering three different robot types. Our results highlight both the potential and current limitations of LMMs in embodied reasoning, pointing towards key challenges and opportunities for future research in robot intelligence. Our data and code will be made publicly available.
title How Good are Foundation Models in Step-by-Step Embodied Reasoning?
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2509.15293