Salvato in:
Dettagli Bibliografici
Autori principali: Cruz, Christian Arzate, Sechayk, Yotam, Igarashi, Takeo, Gomez, Randy
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:https://arxiv.org/abs/2410.00349
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910626896412672
author Cruz, Christian Arzate
Sechayk, Yotam
Igarashi, Takeo
Gomez, Randy
author_facet Cruz, Christian Arzate
Sechayk, Yotam
Igarashi, Takeo
Gomez, Randy
contents Humans use multiple communication channels to interact with each other. For instance, body gestures or facial expressions are commonly used to convey an intent. The use of such non-verbal cues has motivated the development of prediction models. One such approach is predicting arousal and valence (AV) from facial expressions. However, making these models accurate for human-robot interaction (HRI) settings is challenging as it requires handling multiple subjects, challenging conditions, and a wide range of facial expressions. In this paper, we propose a data augmentation (DA) technique to improve the performance of AV predictors using 3D morphable models (3DMM). We then utilize this approach in an HRI setting with a mediator robot and a group of three humans. Our augmentation method creates synthetic sequences for underrepresented values in the AV space of the SEWA dataset, which is the most comprehensive dataset with continuous AV labels. Results show that using our DA method improves the accuracy and robustness of AV prediction in real-time applications. The accuracy of our models on the SEWA dataset is 0.793 for arousal and valence.
format Preprint
id arxiv_https___arxiv_org_abs_2410_00349
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data Augmentation for 3DMM-based Arousal-Valence Prediction for HRI
Cruz, Christian Arzate
Sechayk, Yotam
Igarashi, Takeo
Gomez, Randy
Robotics
Humans use multiple communication channels to interact with each other. For instance, body gestures or facial expressions are commonly used to convey an intent. The use of such non-verbal cues has motivated the development of prediction models. One such approach is predicting arousal and valence (AV) from facial expressions. However, making these models accurate for human-robot interaction (HRI) settings is challenging as it requires handling multiple subjects, challenging conditions, and a wide range of facial expressions. In this paper, we propose a data augmentation (DA) technique to improve the performance of AV predictors using 3D morphable models (3DMM). We then utilize this approach in an HRI setting with a mediator robot and a group of three humans. Our augmentation method creates synthetic sequences for underrepresented values in the AV space of the SEWA dataset, which is the most comprehensive dataset with continuous AV labels. Results show that using our DA method improves the accuracy and robustness of AV prediction in real-time applications. The accuracy of our models on the SEWA dataset is 0.793 for arousal and valence.
title Data Augmentation for 3DMM-based Arousal-Valence Prediction for HRI
topic Robotics
url https://arxiv.org/abs/2410.00349