Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Salehi, Pegah, Sheshkal, Sajad Amouei, Thambawita, Vajira, Riegler, Michael A., Halvorsen, Pål
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912471256662016
author Salehi, Pegah
Sheshkal, Sajad Amouei
Thambawita, Vajira
Riegler, Michael A.
Halvorsen, Pål
author_facet Salehi, Pegah
Sheshkal, Sajad Amouei
Thambawita, Vajira
Riegler, Michael A.
Halvorsen, Pål
contents Dynamic facial emotion is essential for believable AI-generated avatars, yet most systems remain visually static, limiting their use in simulations like virtual training for investigative interviews with abused children. We present a real-time architecture combining Unreal Engine 5 MetaHuman rendering with NVIDIA Omniverse Audio2Face to generate facial expressions from vocal prosody in photorealistic child avatars. Due to limited TTS options, both avatars were voiced using young adult female models from two systems to better fit character profiles, introducing a voice-age mismatch. This confound may affect audiovisual alignment. We used a two-PC setup to decouple speech generation from GPU-intensive rendering, enabling low-latency interaction in desktop and VR. A between-subjects study (N=70) compared audio+visual vs. visual-only conditions as participants rated emotional clarity, facial realism, and empathy for avatars expressing joy, sadness, and anger. While emotions were generally recognized - especially sadness and joy - anger was harder to detect without audio, highlighting the role of voice in high-arousal expressions. Interestingly, silencing clips improved perceived realism by removing mismatches between voice and animation, especially when tone or age felt incongruent. These results emphasize the importance of audiovisual congruence: mismatched voice undermines expression, while a good match can enhance weaker visuals - posing challenges for emotionally coherent avatars in sensitive contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13477
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications
Salehi, Pegah
Sheshkal, Sajad Amouei
Thambawita, Vajira
Riegler, Michael A.
Halvorsen, Pål
Human-Computer Interaction
Computer Vision and Pattern Recognition
68T07, 68U99, 68T45, 91E45
Dynamic facial emotion is essential for believable AI-generated avatars, yet most systems remain visually static, limiting their use in simulations like virtual training for investigative interviews with abused children. We present a real-time architecture combining Unreal Engine 5 MetaHuman rendering with NVIDIA Omniverse Audio2Face to generate facial expressions from vocal prosody in photorealistic child avatars. Due to limited TTS options, both avatars were voiced using young adult female models from two systems to better fit character profiles, introducing a voice-age mismatch. This confound may affect audiovisual alignment. We used a two-PC setup to decouple speech generation from GPU-intensive rendering, enabling low-latency interaction in desktop and VR. A between-subjects study (N=70) compared audio+visual vs. visual-only conditions as participants rated emotional clarity, facial realism, and empathy for avatars expressing joy, sadness, and anger. While emotions were generally recognized - especially sadness and joy - anger was harder to detect without audio, highlighting the role of voice in high-arousal expressions. Interestingly, silencing clips improved perceived realism by removing mismatches between voice and animation, especially when tone or age felt incongruent. These results emphasize the importance of audiovisual congruence: mismatched voice undermines expression, while a good match can enhance weaker visuals - posing challenges for emotionally coherent avatars in sensitive contexts.
title Multimodal Integration Challenges in Emotionally Expressive Child Avatars for Training Applications
topic Human-Computer Interaction
Computer Vision and Pattern Recognition
68T07, 68U99, 68T45, 91E45
url https://arxiv.org/abs/2506.13477