ERR@HRI 2.0 Challenge: Multimodal Detection of Errors and Failures in Human-Robot Conversations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Shiye, Stiber, Maia, Mahmood, Amama, Parreira, Maria Teresa, Ju, Wendy, Spitale, Micol, Gunes, Hatice, Huang, Chien-Ming
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912639078105088
author Cao, Shiye
Stiber, Maia
Mahmood, Amama
Parreira, Maria Teresa
Ju, Wendy
Spitale, Micol
Gunes, Hatice
Huang, Chien-Ming
author_facet Cao, Shiye
Stiber, Maia
Mahmood, Amama
Parreira, Maria Teresa
Ju, Wendy
Spitale, Micol
Gunes, Hatice
Huang, Chien-Ming
contents The integration of large language models (LLMs) into conversational robots has made human-robot conversations more dynamic. Yet, LLM-powered conversational robots remain prone to errors, e.g., misunderstanding user intent, prematurely interrupting users, or failing to respond altogether. Detecting and addressing these failures is critical for preventing conversational breakdowns, avoiding task disruptions, and sustaining user trust. To tackle this problem, the ERR@HRI 2.0 Challenge provides a multimodal dataset of LLM-powered conversational robot failures during human-robot conversations and encourages researchers to benchmark machine learning models designed to detect robot failures. The dataset includes 16 hours of dyadic human-robot interactions, incorporating facial, speech, and head movement features. Each interaction is annotated with the presence or absence of robot errors from the system perspective, and perceived user intention to correct for a mismatch between robot behavior and user expectation. Participants are invited to form teams and develop machine learning models that detect these failures using multimodal data. Submissions will be evaluated using various performance metrics, including detection accuracy and false positive rate. This challenge represents another key step toward improving failure detection in human-robot interaction through social signal analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2507_13468
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ERR@HRI 2.0 Challenge: Multimodal Detection of Errors and Failures in Human-Robot Conversations
Cao, Shiye
Stiber, Maia
Mahmood, Amama
Parreira, Maria Teresa
Ju, Wendy
Spitale, Micol
Gunes, Hatice
Huang, Chien-Ming
Robotics
Artificial Intelligence
Human-Computer Interaction
The integration of large language models (LLMs) into conversational robots has made human-robot conversations more dynamic. Yet, LLM-powered conversational robots remain prone to errors, e.g., misunderstanding user intent, prematurely interrupting users, or failing to respond altogether. Detecting and addressing these failures is critical for preventing conversational breakdowns, avoiding task disruptions, and sustaining user trust. To tackle this problem, the ERR@HRI 2.0 Challenge provides a multimodal dataset of LLM-powered conversational robot failures during human-robot conversations and encourages researchers to benchmark machine learning models designed to detect robot failures. The dataset includes 16 hours of dyadic human-robot interactions, incorporating facial, speech, and head movement features. Each interaction is annotated with the presence or absence of robot errors from the system perspective, and perceived user intention to correct for a mismatch between robot behavior and user expectation. Participants are invited to form teams and develop machine learning models that detect these failures using multimodal data. Submissions will be evaluated using various performance metrics, including detection accuracy and false positive rate. This challenge represents another key step toward improving failure detection in human-robot interaction through social signal analysis.
title ERR@HRI 2.0 Challenge: Multimodal Detection of Errors and Failures in Human-Robot Conversations
topic Robotics
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2507.13468