Enhancing XR Auditory Realism via Multimodal Scene-Aware Acoustic Rendering

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Tianyu, Li, Jihan, Zu, Penghe, Sahay, Pranav, Kim, Maruchi, Obeng-Marnu, Jack, Miller, Farley, Qian, Xun, Passarella, Katrina, Rachumalla, Mahitha, Nongpiur, Rajeev, Shin, D.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914159080243200
author Xu, Tianyu
Li, Jihan
Zu, Penghe
Sahay, Pranav
Kim, Maruchi
Obeng-Marnu, Jack
Miller, Farley
Qian, Xun
Passarella, Katrina
Rachumalla, Mahitha
Nongpiur, Rajeev
Shin, D.
author_facet Xu, Tianyu
Li, Jihan
Zu, Penghe
Sahay, Pranav
Kim, Maruchi
Obeng-Marnu, Jack
Miller, Farley
Qian, Xun
Passarella, Katrina
Rachumalla, Mahitha
Nongpiur, Rajeev
Shin, D.
contents In Extended Reality (XR), rendering sound that accurately simulates real-world acoustics is pivotal in creating lifelike and believable virtual experiences. However, existing XR spatial audio rendering methods often struggle with real-time adaptation to diverse physical scenes, causing a sensory mismatch between visual and auditory cues that disrupts user immersion. To address this, we introduce SAMOSA, a novel on-device system that renders spatially accurate sound by dynamically adapting to its physical environment. SAMOSA leverages a synergistic multimodal scene representation by fusing real-time estimations of room geometry, surface materials, and semantic-driven acoustic context. This rich representation then enables efficient acoustic calibration via scene priors, allowing the system to synthesize a highly realistic Room Impulse Response (RIR). We validate our system through technical evaluation using acoustic metrics for RIR synthesis across various room configurations and sound types, alongside an expert evaluation (N=12). Evaluation results demonstrate SAMOSA's feasibility and efficacy in enhancing XR auditory realism.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11930
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing XR Auditory Realism via Multimodal Scene-Aware Acoustic Rendering
Xu, Tianyu
Li, Jihan
Zu, Penghe
Sahay, Pranav
Kim, Maruchi
Obeng-Marnu, Jack
Miller, Farley
Qian, Xun
Passarella, Katrina
Rachumalla, Mahitha
Nongpiur, Rajeev
Shin, D.
Human-Computer Interaction
Computer Vision and Pattern Recognition
Machine Learning
Sound
H.5.1; H.5.5; I.2.7; I.2.10
In Extended Reality (XR), rendering sound that accurately simulates real-world acoustics is pivotal in creating lifelike and believable virtual experiences. However, existing XR spatial audio rendering methods often struggle with real-time adaptation to diverse physical scenes, causing a sensory mismatch between visual and auditory cues that disrupts user immersion. To address this, we introduce SAMOSA, a novel on-device system that renders spatially accurate sound by dynamically adapting to its physical environment. SAMOSA leverages a synergistic multimodal scene representation by fusing real-time estimations of room geometry, surface materials, and semantic-driven acoustic context. This rich representation then enables efficient acoustic calibration via scene priors, allowing the system to synthesize a highly realistic Room Impulse Response (RIR). We validate our system through technical evaluation using acoustic metrics for RIR synthesis across various room configurations and sound types, alongside an expert evaluation (N=12). Evaluation results demonstrate SAMOSA's feasibility and efficacy in enhancing XR auditory realism.
title Enhancing XR Auditory Realism via Multimodal Scene-Aware Acoustic Rendering
topic Human-Computer Interaction
Computer Vision and Pattern Recognition
Machine Learning
Sound
H.5.1; H.5.5; I.2.7; I.2.10
url https://arxiv.org/abs/2511.11930