Novel-View Acoustic Synthesis from 3D Reconstructed Rooms

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ahn, Byeongjoo, Yang, Karren, Hamilton, Brian, Sheaffer, Jonathan, Ranjan, Anurag, Sarabia, Miguel, Tuzel, Oncel, Chang, Jen-Hao Rick
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913468837265408
author Ahn, Byeongjoo
Yang, Karren
Hamilton, Brian
Sheaffer, Jonathan
Ranjan, Anurag
Sarabia, Miguel
Tuzel, Oncel
Chang, Jen-Hao Rick
author_facet Ahn, Byeongjoo
Yang, Karren
Hamilton, Brian
Sheaffer, Jonathan
Ranjan, Anurag
Sarabia, Miguel
Tuzel, Oncel
Chang, Jen-Hao Rick
contents We investigate the benefit of combining blind audio recordings with 3D scene information for novel-view acoustic synthesis. Given audio recordings from 2-4 microphones and the 3D geometry and material of a scene containing multiple unknown sound sources, we estimate the sound anywhere in the scene. We identify the main challenges of novel-view acoustic synthesis as sound source localization, separation, and dereverberation. While naively training an end-to-end network fails to produce high-quality results, we show that incorporating room impulse responses (RIRs) derived from 3D reconstructed rooms enables the same network to jointly tackle these tasks. Our method outperforms existing methods designed for the individual tasks, demonstrating its effectiveness at utilizing 3D visual information. In a simulated study on the Matterport3D-NVAS dataset, our model achieves near-perfect accuracy on source localization, a PSNR of 26.44dB and a SDR of 14.23dB for source separation and dereverberation, resulting in a PSNR of 25.55 dB and a SDR of 14.20 dB on novel-view acoustic synthesis. We release our code and model on our project website at https://github.com/apple/ml-nvas3d. Please wear headphones when listening to the results.
format Preprint
id arxiv_https___arxiv_org_abs_2310_15130
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Novel-View Acoustic Synthesis from 3D Reconstructed Rooms
Ahn, Byeongjoo
Yang, Karren
Hamilton, Brian
Sheaffer, Jonathan
Ranjan, Anurag
Sarabia, Miguel
Tuzel, Oncel
Chang, Jen-Hao Rick
Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
We investigate the benefit of combining blind audio recordings with 3D scene information for novel-view acoustic synthesis. Given audio recordings from 2-4 microphones and the 3D geometry and material of a scene containing multiple unknown sound sources, we estimate the sound anywhere in the scene. We identify the main challenges of novel-view acoustic synthesis as sound source localization, separation, and dereverberation. While naively training an end-to-end network fails to produce high-quality results, we show that incorporating room impulse responses (RIRs) derived from 3D reconstructed rooms enables the same network to jointly tackle these tasks. Our method outperforms existing methods designed for the individual tasks, demonstrating its effectiveness at utilizing 3D visual information. In a simulated study on the Matterport3D-NVAS dataset, our model achieves near-perfect accuracy on source localization, a PSNR of 26.44dB and a SDR of 14.23dB for source separation and dereverberation, resulting in a PSNR of 25.55 dB and a SDR of 14.20 dB on novel-view acoustic synthesis. We release our code and model on our project website at https://github.com/apple/ml-nvas3d. Please wear headphones when listening to the results.
title Novel-View Acoustic Synthesis from 3D Reconstructed Rooms
topic Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
url https://arxiv.org/abs/2310.15130