Spatial Audio Rendering for Real-Time Speech Translation in Virtual Meetings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Geleta, Margarita, Sodoma, Hong, Gamper, Hannes
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909899589419008
author Geleta, Margarita
Sodoma, Hong
Gamper, Hannes
author_facet Geleta, Margarita
Sodoma, Hong
Gamper, Hannes
contents Language barriers in virtual meetings remain a persistent challenge to global collaboration. Real-time translation offers promise, yet current integrations often neglect perceptual cues. This study investigates how spatial audio rendering of translated speech influences comprehension, cognitive load, and user experience in multilingual meetings. We conducted a within-subjects experiment with 8 bilingual confederates and 47 participants simulating global team meetings with English translations of Greek, Kannada, Mandarin Chinese, and Ukrainian - languages selected for their diversity in grammar, script, and resource availability. Participants experienced four audio conditions: spatial audio with and without background reverberation, and two non-spatial configurations (diotic, monaural). We measured listener comprehension accuracy, workload ratings, satisfaction scores, and qualitative feedback. Spatially-rendered translations doubled comprehension compared to non-spatial audio. Participants reported greater clarity and engagement when spatial cues and voice timbre differentiation were present. We discuss design implications for integrating real-time translation into meeting platforms, advancing inclusive, cross-language communication in telepresence systems.
format Preprint
id arxiv_https___arxiv_org_abs_2511_09525
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Spatial Audio Rendering for Real-Time Speech Translation in Virtual Meetings
Geleta, Margarita
Sodoma, Hong
Gamper, Hannes
Human-Computer Interaction
Sound
H.5; H.5.5
Language barriers in virtual meetings remain a persistent challenge to global collaboration. Real-time translation offers promise, yet current integrations often neglect perceptual cues. This study investigates how spatial audio rendering of translated speech influences comprehension, cognitive load, and user experience in multilingual meetings. We conducted a within-subjects experiment with 8 bilingual confederates and 47 participants simulating global team meetings with English translations of Greek, Kannada, Mandarin Chinese, and Ukrainian - languages selected for their diversity in grammar, script, and resource availability. Participants experienced four audio conditions: spatial audio with and without background reverberation, and two non-spatial configurations (diotic, monaural). We measured listener comprehension accuracy, workload ratings, satisfaction scores, and qualitative feedback. Spatially-rendered translations doubled comprehension compared to non-spatial audio. Participants reported greater clarity and engagement when spatial cues and voice timbre differentiation were present. We discuss design implications for integrating real-time translation into meeting platforms, advancing inclusive, cross-language communication in telepresence systems.
title Spatial Audio Rendering for Real-Time Speech Translation in Virtual Meetings
topic Human-Computer Interaction
Sound
H.5; H.5.5
url https://arxiv.org/abs/2511.09525