Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fan, Congyi, Guan, Jian, Lin, Youtian, Xu, Dongli, Ye, Tong, Zhu, Qiaoxi, Feng, Pengming, Wang, Wenwu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911402639228928
author Fan, Congyi
Guan, Jian
Lin, Youtian
Xu, Dongli
Ye, Tong
Zhu, Qiaoxi
Feng, Pengming
Wang, Wenwu
author_facet Fan, Congyi
Guan, Jian
Lin, Youtian
Xu, Dongli
Ye, Tong
Zhu, Qiaoxi
Feng, Pengming
Wang, Wenwu
contents Spatial audio is essential for immersive experiences, yet novel-view acoustic synthesis (NVAS) remains challenging due to complex physical phenomena such as reflection, diffraction, and material absorption. Existing methods based on single-view or panoramic inputs improve spatial fidelity but fail to capture global geometry and semantic cues such as object layout and material properties. To address this, we propose Phys-NVAS, the first physics-aware NVAS framework that integrates spatial geometry modeling with vision-language semantic priors. A global 3D acoustic environment is reconstructed from multi-view images and depth maps to estimate room size and shape, enhancing spatial awareness of sound propagation. Meanwhile, a vision-language model extracts physics-aware priors of objects, layouts, and materials, capturing absorption and reflection beyond geometry. An acoustic feature fusion adapter unifies these cues into a physics-aware representation for binaural generation. Experiments on RWAVS demonstrate that Phys-NVAS yields binaural audio with improved realism and physical consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19712
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling
Fan, Congyi
Guan, Jian
Lin, Youtian
Xu, Dongli
Ye, Tong
Zhu, Qiaoxi
Feng, Pengming
Wang, Wenwu
Sound
Multimedia
Spatial audio is essential for immersive experiences, yet novel-view acoustic synthesis (NVAS) remains challenging due to complex physical phenomena such as reflection, diffraction, and material absorption. Existing methods based on single-view or panoramic inputs improve spatial fidelity but fail to capture global geometry and semantic cues such as object layout and material properties. To address this, we propose Phys-NVAS, the first physics-aware NVAS framework that integrates spatial geometry modeling with vision-language semantic priors. A global 3D acoustic environment is reconstructed from multi-view images and depth maps to estimate room size and shape, enhancing spatial awareness of sound propagation. Meanwhile, a vision-language model extracts physics-aware priors of objects, layouts, and materials, capturing absorption and reflection beyond geometry. An acoustic feature fusion adapter unifies these cues into a physics-aware representation for binaural generation. Experiments on RWAVS demonstrate that Phys-NVAS yields binaural audio with improved realism and physical consistency.
title Physics-Aware Novel-View Acoustic Synthesis with Vision-Language Priors and 3D Acoustic Environment Modeling
topic Sound
Multimedia
url https://arxiv.org/abs/2601.19712