Saved in:
Bibliographic Details
Main Authors: Bečková, Iveta, Pócoš, Štefan, Belgiovine, Giulia, Matarese, Marco, Eldardeer, Omar, Sciutti, Alessandra, Mazzola, Carlo
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2407.03340
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909470441865216
author Bečková, Iveta
Pócoš, Štefan
Belgiovine, Giulia
Matarese, Marco
Eldardeer, Omar
Sciutti, Alessandra
Mazzola, Carlo
author_facet Bečková, Iveta
Pócoš, Štefan
Belgiovine, Giulia
Matarese, Marco
Eldardeer, Omar
Sciutti, Alessandra
Mazzola, Carlo
contents The addressee estimation (understanding to whom somebody is talking) is a fundamental task for human activity recognition in multi-party conversation scenarios. Specifically, in the field of human-robot interaction, it becomes even more crucial to enable social robots to participate in such interactive contexts. However, it is usually implemented as a binary classification task, restricting the robot's capability to estimate whether it was addressed \review{or not, which} limits its interactive skills. For a social robot to gain the trust of humans, it is also important to manifest a certain level of transparency and explainability. Explainable artificial intelligence thus plays a significant role in the current machine learning applications and models, to provide explanations for their decisions besides excellent performance. In our work, we a) present an addressee estimation model with improved performance in comparison with the previous state-of-the-art; b) further modify this model to include inherently explainable attention-based segments; c) implement the explainable addressee estimation as part of a modular cognitive architecture for multi-party conversation in an iCub robot; d) validate the real-time performance of the explainable model in multi-party human-robot interaction; e) propose several ways to incorporate explainability and transparency in the aforementioned architecture; and f) perform an online user study to analyze the effect of various explanations on how human participants perceive the robot.
format Preprint
id arxiv_https___arxiv_org_abs_2407_03340
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Multi-Modal Explainability Approach for Human-Aware Robots in Multi-Party Conversation
Bečková, Iveta
Pócoš, Štefan
Belgiovine, Giulia
Matarese, Marco
Eldardeer, Omar
Sciutti, Alessandra
Mazzola, Carlo
Artificial Intelligence
Computation and Language
Human-Computer Interaction
Machine Learning
Robotics
Image and Video Processing
I.4.8; I.2.10; I.2.9; I.2.11; J.4
The addressee estimation (understanding to whom somebody is talking) is a fundamental task for human activity recognition in multi-party conversation scenarios. Specifically, in the field of human-robot interaction, it becomes even more crucial to enable social robots to participate in such interactive contexts. However, it is usually implemented as a binary classification task, restricting the robot's capability to estimate whether it was addressed \review{or not, which} limits its interactive skills. For a social robot to gain the trust of humans, it is also important to manifest a certain level of transparency and explainability. Explainable artificial intelligence thus plays a significant role in the current machine learning applications and models, to provide explanations for their decisions besides excellent performance. In our work, we a) present an addressee estimation model with improved performance in comparison with the previous state-of-the-art; b) further modify this model to include inherently explainable attention-based segments; c) implement the explainable addressee estimation as part of a modular cognitive architecture for multi-party conversation in an iCub robot; d) validate the real-time performance of the explainable model in multi-party human-robot interaction; e) propose several ways to incorporate explainability and transparency in the aforementioned architecture; and f) perform an online user study to analyze the effect of various explanations on how human participants perceive the robot.
title A Multi-Modal Explainability Approach for Human-Aware Robots in Multi-Party Conversation
topic Artificial Intelligence
Computation and Language
Human-Computer Interaction
Machine Learning
Robotics
Image and Video Processing
I.4.8; I.2.10; I.2.9; I.2.11; J.4
url https://arxiv.org/abs/2407.03340