Blind Spatial Impulse Response Generation from Separate Room- and Scene-Specific Information

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lluís, Francesc, Meyer-Kahlen, Nils
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910617467617280
author Lluís, Francesc
Meyer-Kahlen, Nils
author_facet Lluís, Francesc
Meyer-Kahlen, Nils
contents For audio in augmented reality (AR), knowledge of the users' real acoustic environment is crucial for rendering virtual sounds that seamlessly blend into the environment. As acoustic measurements are usually not feasible in practical AR applications, information about the room needs to be inferred from available sound sources. Then, additional sound sources can be rendered with the same room acoustic qualities. Crucially, these are placed at different positions than the sources available for estimation. Here, we propose to use an encoder network trained using a contrastive loss that maps input sounds to a low-dimensional feature space representing only room-specific information. Then, a diffusion-based spatial room impulse response generator is trained to take the latent space and generate a new response, given a new source-receiver position. We show how both room- and position-specific parameters are considered in the final output.
format Preprint
id arxiv_https___arxiv_org_abs_2409_14971
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Blind Spatial Impulse Response Generation from Separate Room- and Scene-Specific Information
Lluís, Francesc
Meyer-Kahlen, Nils
Sound
Machine Learning
Audio and Speech Processing
For audio in augmented reality (AR), knowledge of the users' real acoustic environment is crucial for rendering virtual sounds that seamlessly blend into the environment. As acoustic measurements are usually not feasible in practical AR applications, information about the room needs to be inferred from available sound sources. Then, additional sound sources can be rendered with the same room acoustic qualities. Crucially, these are placed at different positions than the sources available for estimation. Here, we propose to use an encoder network trained using a contrastive loss that maps input sounds to a low-dimensional feature space representing only room-specific information. Then, a diffusion-based spatial room impulse response generator is trained to take the latent space and generate a new response, given a new source-receiver position. We show how both room- and position-specific parameters are considered in the final output.
title Blind Spatial Impulse Response Generation from Separate Room- and Scene-Specific Information
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2409.14971