Multimodal Deep Learning Method for Real-Time Spatial Room Impulse Response Computing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhiyu, Yue, Xinwen, Zhao, Shenghui, Wang, Jing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911571816480768
author Li, Zhiyu
Yue, Xinwen
Zhao, Shenghui
Wang, Jing
author_facet Li, Zhiyu
Yue, Xinwen
Zhao, Shenghui
Wang, Jing
contents We propose a multimodal deep learning model for VR auralization that generates spatial room impulse responses (SRIRs) in real time to reconstruct scene-specific auditory perception. Employing SRIRs as the output reduces computational complexity and facilitates integration with personalized head-related transfer functions. The model takes two modalities as input: scene information and waveforms, where the waveform corresponds to the low-order reflections (LoR). LoR can be efficiently computed using geometrical acoustics (GA) but remains difficult for deep learning models to predict accurately. Scene geometry, acoustic properties, source coordinates, and listener coordinates are first used to compute LoR in real time via GA, and both LoR and these features are subsequently provided as inputs to the model. A new dataset was constructed, consisting of multiple scenes and their corresponding SRIRs. The dataset exhibits greater diversity. Experimental results demonstrate the superior performance of the proposed model.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05545
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multimodal Deep Learning Method for Real-Time Spatial Room Impulse Response Computing
Li, Zhiyu
Yue, Xinwen
Zhao, Shenghui
Wang, Jing
Audio and Speech Processing
We propose a multimodal deep learning model for VR auralization that generates spatial room impulse responses (SRIRs) in real time to reconstruct scene-specific auditory perception. Employing SRIRs as the output reduces computational complexity and facilitates integration with personalized head-related transfer functions. The model takes two modalities as input: scene information and waveforms, where the waveform corresponds to the low-order reflections (LoR). LoR can be efficiently computed using geometrical acoustics (GA) but remains difficult for deep learning models to predict accurately. Scene geometry, acoustic properties, source coordinates, and listener coordinates are first used to compute LoR in real time via GA, and both LoR and these features are subsequently provided as inputs to the model. A new dataset was constructed, consisting of multiple scenes and their corresponding SRIRs. The dataset exhibits greater diversity. Experimental results demonstrate the superior performance of the proposed model.
title Multimodal Deep Learning Method for Real-Time Spatial Room Impulse Response Computing
topic Audio and Speech Processing
url https://arxiv.org/abs/2604.05545