Pragmatic Embodied Spoken Instruction Following in Human-Robot Collaboration with Theory of Mind

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ying, Lance, Li, Xinyi, Aarya, Shivam, Fang, Yizirui, Yin, Yifan, Liu, Jason Xinyu, Tellex, Stefanie, Tenenbaum, Joshua B., Shu, Tianmin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916991139315712
author Ying, Lance
Li, Xinyi
Aarya, Shivam
Fang, Yizirui
Yin, Yifan
Liu, Jason Xinyu
Tellex, Stefanie
Tenenbaum, Joshua B.
Shu, Tianmin
author_facet Ying, Lance
Li, Xinyi
Aarya, Shivam
Fang, Yizirui
Yin, Yifan
Liu, Jason Xinyu
Tellex, Stefanie
Tenenbaum, Joshua B.
Shu, Tianmin
contents Spoken language instructions are ubiquitous in agent collaboration. However, in real-world human-robot collaboration, following human spoken instructions can be challenging due to various speaker and environmental factors, such as background noise or mispronunciation. When faced with noisy auditory inputs, humans can leverage the collaborative context in the embodied environment to interpret noisy spoken instructions and take pragmatic assistive actions. In this paper, we present a cognitively inspired neurosymbolic model, Spoken Instruction Following through Theory of Mind (SIFToM), which leverages a Vision-Language Model with model-based mental inference to enable robots to pragmatically follow human instructions under diverse speech conditions. We test SIFToM in both simulated environments (VirtualHome) and real-world human-robot collaborative settings with human evaluations. Results show that SIFToM can significantly improve the performance of a lightweight base VLM (Gemini 2.5 Flash), outperforming state-of-the-art VLMs (Gemini 2.5 Pro) and approaching human-level accuracy on challenging spoken instruction following tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2409_10849
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Pragmatic Embodied Spoken Instruction Following in Human-Robot Collaboration with Theory of Mind
Ying, Lance
Li, Xinyi
Aarya, Shivam
Fang, Yizirui
Yin, Yifan
Liu, Jason Xinyu
Tellex, Stefanie
Tenenbaum, Joshua B.
Shu, Tianmin
Robotics
Artificial Intelligence
Human-Computer Interaction
Multiagent Systems
Spoken language instructions are ubiquitous in agent collaboration. However, in real-world human-robot collaboration, following human spoken instructions can be challenging due to various speaker and environmental factors, such as background noise or mispronunciation. When faced with noisy auditory inputs, humans can leverage the collaborative context in the embodied environment to interpret noisy spoken instructions and take pragmatic assistive actions. In this paper, we present a cognitively inspired neurosymbolic model, Spoken Instruction Following through Theory of Mind (SIFToM), which leverages a Vision-Language Model with model-based mental inference to enable robots to pragmatically follow human instructions under diverse speech conditions. We test SIFToM in both simulated environments (VirtualHome) and real-world human-robot collaborative settings with human evaluations. Results show that SIFToM can significantly improve the performance of a lightweight base VLM (Gemini 2.5 Flash), outperforming state-of-the-art VLMs (Gemini 2.5 Pro) and approaching human-level accuracy on challenging spoken instruction following tasks.
title Pragmatic Embodied Spoken Instruction Following in Human-Robot Collaboration with Theory of Mind
topic Robotics
Artificial Intelligence
Human-Computer Interaction
Multiagent Systems
url https://arxiv.org/abs/2409.10849