Multimodal Human-Autonomous Agents Interaction Using Pre-Trained Language and Visual Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nwankwo, Linus, Rueckert, Elmar
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913628269051904
author Nwankwo, Linus
Rueckert, Elmar
author_facet Nwankwo, Linus
Rueckert, Elmar
contents In this paper, we extended the method proposed in [21] to enable humans to interact naturally with autonomous agents through vocal and textual conversations. Our extended method exploits the inherent capabilities of pre-trained large language models (LLMs), multimodal visual language models (VLMs), and speech recognition (SR) models to decode the high-level natural language conversations and semantic understanding of the robot's task environment, and abstract them to the robot's actionable commands or queries. We performed a quantitative evaluation of our framework's natural vocal conversation understanding with participants from different racial backgrounds and English language accents. The participants interacted with the robot using both spoken and textual instructional commands. Based on the logged interaction data, our framework achieved 87.55% vocal commands decoding accuracy, 86.27% commands execution success, and an average latency of 0.89 seconds from receiving the participants' vocal chat commands to initiating the robot's actual physical action. The video demonstrations of this paper can be found at https://linusnep.github.io/MTCC-IRoNL/.
format Preprint
id arxiv_https___arxiv_org_abs_2403_12273
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multimodal Human-Autonomous Agents Interaction Using Pre-Trained Language and Visual Foundation Models
Nwankwo, Linus
Rueckert, Elmar
Robotics
In this paper, we extended the method proposed in [21] to enable humans to interact naturally with autonomous agents through vocal and textual conversations. Our extended method exploits the inherent capabilities of pre-trained large language models (LLMs), multimodal visual language models (VLMs), and speech recognition (SR) models to decode the high-level natural language conversations and semantic understanding of the robot's task environment, and abstract them to the robot's actionable commands or queries. We performed a quantitative evaluation of our framework's natural vocal conversation understanding with participants from different racial backgrounds and English language accents. The participants interacted with the robot using both spoken and textual instructional commands. Based on the logged interaction data, our framework achieved 87.55% vocal commands decoding accuracy, 86.27% commands execution success, and an average latency of 0.89 seconds from receiving the participants' vocal chat commands to initiating the robot's actual physical action. The video demonstrations of this paper can be found at https://linusnep.github.io/MTCC-IRoNL/.
title Multimodal Human-Autonomous Agents Interaction Using Pre-Trained Language and Visual Foundation Models
topic Robotics
url https://arxiv.org/abs/2403.12273