NeRAF: 3D Scene Infused Neural Radiance and Acoustic Fields

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Brunetto, Amandine, Hornauer, Sascha, Moutarde, Fabien
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915529350971392
author Brunetto, Amandine
Hornauer, Sascha
Moutarde, Fabien
author_facet Brunetto, Amandine
Hornauer, Sascha
Moutarde, Fabien
contents Sound plays a major role in human perception. Along with vision, it provides essential information for understanding our surroundings. Despite advances in neural implicit representations, learning acoustics that align with visual scenes remains a challenge. We propose NeRAF, a method that jointly learns acoustic and radiance fields. NeRAF synthesizes both novel views and spatialized room impulse responses (RIR) at new positions by conditioning the acoustic field on 3D scene geometric and appearance priors from the radiance field. The generated RIR can be applied to auralize any audio signal. Each modality can be rendered independently and at spatially distinct positions, offering greater versatility. We demonstrate that NeRAF generates high-quality audio on SoundSpaces and RAF datasets, achieving significant performance improvements over prior methods while being more data-efficient. Additionally, NeRAF enhances novel view synthesis of complex scenes trained with sparse data through cross-modal learning. NeRAF is designed as a Nerfstudio module, providing convenient access to realistic audio-visual generation.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18213
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NeRAF: 3D Scene Infused Neural Radiance and Acoustic Fields
Brunetto, Amandine
Hornauer, Sascha
Moutarde, Fabien
Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
Sound plays a major role in human perception. Along with vision, it provides essential information for understanding our surroundings. Despite advances in neural implicit representations, learning acoustics that align with visual scenes remains a challenge. We propose NeRAF, a method that jointly learns acoustic and radiance fields. NeRAF synthesizes both novel views and spatialized room impulse responses (RIR) at new positions by conditioning the acoustic field on 3D scene geometric and appearance priors from the radiance field. The generated RIR can be applied to auralize any audio signal. Each modality can be rendered independently and at spatially distinct positions, offering greater versatility. We demonstrate that NeRAF generates high-quality audio on SoundSpaces and RAF datasets, achieving significant performance improvements over prior methods while being more data-efficient. Additionally, NeRAF enhances novel view synthesis of complex scenes trained with sparse data through cross-modal learning. NeRAF is designed as a Nerfstudio module, providing convenient access to realistic audio-visual generation.
title NeRAF: 3D Scene Infused Neural Radiance and Acoustic Fields
topic Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
url https://arxiv.org/abs/2405.18213