Agreeing to Interact in Human-Robot Interaction using Large Language Models and Vision Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sasabuchi, Kazuhiro, Wake, Naoki, Kanehira, Atsushi, Takamatsu, Jun, Ikeuchi, Katsushi
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908275857948672
author Sasabuchi, Kazuhiro
Wake, Naoki
Kanehira, Atsushi
Takamatsu, Jun
Ikeuchi, Katsushi
author_facet Sasabuchi, Kazuhiro
Wake, Naoki
Kanehira, Atsushi
Takamatsu, Jun
Ikeuchi, Katsushi
contents In human-robot interaction (HRI), the beginning of an interaction is often complex. Whether the robot should communicate with the human is dependent on several situational factors (e.g., the current human's activity, urgency of the interaction, etc.). We test whether large language models (LLM) and vision language models (VLM) can provide solutions to this problem. We compare four different system-design patterns using LLMs and VLMs, and test on a test set containing 84 human-robot situations. The test set mixes several publicly available datasets and also includes situations where the appropriate action to take is open-ended. Our results using the GPT-4o and Phi-3 Vision model indicate that LLMs and VLMs are capable of handling interaction beginnings when the desired actions are clear, however, challenge remains in the open-ended situations where the model must balance between the human and robot situation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15491
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Agreeing to Interact in Human-Robot Interaction using Large Language Models and Vision Language Models
Sasabuchi, Kazuhiro
Wake, Naoki
Kanehira, Atsushi
Takamatsu, Jun
Ikeuchi, Katsushi
Human-Computer Interaction
Computation and Language
Machine Learning
Robotics
In human-robot interaction (HRI), the beginning of an interaction is often complex. Whether the robot should communicate with the human is dependent on several situational factors (e.g., the current human's activity, urgency of the interaction, etc.). We test whether large language models (LLM) and vision language models (VLM) can provide solutions to this problem. We compare four different system-design patterns using LLMs and VLMs, and test on a test set containing 84 human-robot situations. The test set mixes several publicly available datasets and also includes situations where the appropriate action to take is open-ended. Our results using the GPT-4o and Phi-3 Vision model indicate that LLMs and VLMs are capable of handling interaction beginnings when the desired actions are clear, however, challenge remains in the open-ended situations where the model must balance between the human and robot situation.
title Agreeing to Interact in Human-Robot Interaction using Large Language Models and Vision Language Models
topic Human-Computer Interaction
Computation and Language
Machine Learning
Robotics
url https://arxiv.org/abs/2503.15491