Multimodal Large Language Models for Real-Time Situated Reasoning

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Abbo, Giulio Antonio, Lenaerts, Senne, Belpaeme, Tony
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910008648663040
author Abbo, Giulio Antonio
Lenaerts, Senne
Belpaeme, Tony
author_facet Abbo, Giulio Antonio
Lenaerts, Senne
Belpaeme, Tony
contents In this work, we explore how multimodal large language models can support real-time context- and value-aware decision-making. To do so, we combine the GPT-4o language model with a TurtleBot 4 platform simulating a smart vacuum cleaning robot in a home. The model evaluates the environment through vision input and determines whether it is appropriate to initiate cleaning. The system highlights the ability of these models to reason about domestic activities, social norms, and user preferences and take nuanced decisions aligned with the values of the people involved, such as cleanliness, comfort, and safety. We demonstrate the system in a realistic home environment, showing its ability to infer context and values from limited visual input. Our results highlight the promise of multimodal large language models in enhancing robotic autonomy and situational awareness, while also underscoring challenges related to consistency, bias, and real-time performance.
format Preprint
id arxiv_https___arxiv_org_abs_2602_01880
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multimodal Large Language Models for Real-Time Situated Reasoning
Abbo, Giulio Antonio
Lenaerts, Senne
Belpaeme, Tony
Robotics
In this work, we explore how multimodal large language models can support real-time context- and value-aware decision-making. To do so, we combine the GPT-4o language model with a TurtleBot 4 platform simulating a smart vacuum cleaning robot in a home. The model evaluates the environment through vision input and determines whether it is appropriate to initiate cleaning. The system highlights the ability of these models to reason about domestic activities, social norms, and user preferences and take nuanced decisions aligned with the values of the people involved, such as cleanliness, comfort, and safety. We demonstrate the system in a realistic home environment, showing its ability to infer context and values from limited visual input. Our results highlight the promise of multimodal large language models in enhancing robotic autonomy and situational awareness, while also underscoring challenges related to consistency, bias, and real-time performance.
title Multimodal Large Language Models for Real-Time Situated Reasoning
topic Robotics
url https://arxiv.org/abs/2602.01880