Enhancing Robot Explanation Capabilities through Vision-Language Models: a Preliminary Study by Interpreting Visual Inputs for Improved Human-Robot Interaction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sobrín-Hidalgo, David, González-Santamarta, Miguel Ángel, Guerrero-Higueras, Ángel Manuel, Rodríguez-Lera, Francisco Javier, Matellán-Olivera, Vicente
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913315472539648
author Sobrín-Hidalgo, David
González-Santamarta, Miguel Ángel
Guerrero-Higueras, Ángel Manuel
Rodríguez-Lera, Francisco Javier
Matellán-Olivera, Vicente
author_facet Sobrín-Hidalgo, David
González-Santamarta, Miguel Ángel
Guerrero-Higueras, Ángel Manuel
Rodríguez-Lera, Francisco Javier
Matellán-Olivera, Vicente
contents This paper presents an improved system based on our prior work, designed to create explanations for autonomous robot actions during Human-Robot Interaction (HRI). Previously, we developed a system that used Large Language Models (LLMs) to interpret logs and produce natural language explanations. In this study, we expand our approach by incorporating Vision-Language Models (VLMs), enabling the system to analyze textual logs with the added context of visual input. This method allows for generating explanations that combine data from the robot's logs and the images it captures. We tested this enhanced system on a basic navigation task where the robot needs to avoid a human obstacle. The findings from this preliminary study indicate that adding visual interpretation improves our system's explanations by precisely identifying obstacles and increasing the accuracy of the explanations provided.
format Preprint
id arxiv_https___arxiv_org_abs_2404_09705
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Robot Explanation Capabilities through Vision-Language Models: a Preliminary Study by Interpreting Visual Inputs for Improved Human-Robot Interaction
Sobrín-Hidalgo, David
González-Santamarta, Miguel Ángel
Guerrero-Higueras, Ángel Manuel
Rodríguez-Lera, Francisco Javier
Matellán-Olivera, Vicente
Robotics
This paper presents an improved system based on our prior work, designed to create explanations for autonomous robot actions during Human-Robot Interaction (HRI). Previously, we developed a system that used Large Language Models (LLMs) to interpret logs and produce natural language explanations. In this study, we expand our approach by incorporating Vision-Language Models (VLMs), enabling the system to analyze textual logs with the added context of visual input. This method allows for generating explanations that combine data from the robot's logs and the images it captures. We tested this enhanced system on a basic navigation task where the robot needs to avoid a human obstacle. The findings from this preliminary study indicate that adding visual interpretation improves our system's explanations by precisely identifying obstacles and increasing the accuracy of the explanations provided.
title Enhancing Robot Explanation Capabilities through Vision-Language Models: a Preliminary Study by Interpreting Visual Inputs for Improved Human-Robot Interaction
topic Robotics
url https://arxiv.org/abs/2404.09705