Visual Acoustic Fields

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Yuelei, Kim, Hyunjin, Zhan, Fangneng, Qiu, Ri-Zhao, Ji, Mazeyu, Shan, Xiaojun, Zou, Xueyan, Liang, Paul, Pfister, Hanspeter, Wang, Xiaolong
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912302147567616
author Li, Yuelei
Kim, Hyunjin
Zhan, Fangneng
Qiu, Ri-Zhao
Ji, Mazeyu
Shan, Xiaojun
Zou, Xueyan
Liang, Paul
Pfister, Hanspeter
Wang, Xiaolong
author_facet Li, Yuelei
Kim, Hyunjin
Zhan, Fangneng
Qiu, Ri-Zhao
Ji, Mazeyu
Shan, Xiaojun
Zou, Xueyan
Liang, Paul
Pfister, Hanspeter
Wang, Xiaolong
contents Objects produce different sounds when hit, and humans can intuitively infer how an object might sound based on its appearance and material properties. Inspired by this intuition, we propose Visual Acoustic Fields, a framework that bridges hitting sounds and visual signals within a 3D space using 3D Gaussian Splatting (3DGS). Our approach features two key modules: sound generation and sound localization. The sound generation module leverages a conditional diffusion model, which takes multiscale features rendered from a feature-augmented 3DGS to generate realistic hitting sounds. Meanwhile, the sound localization module enables querying the 3D scene, represented by the feature-augmented 3DGS, to localize hitting positions based on the sound sources. To support this framework, we introduce a novel pipeline for collecting scene-level visual-sound sample pairs, achieving alignment between captured images, impact locations, and corresponding sounds. To the best of our knowledge, this is the first dataset to connect visual and acoustic signals in a 3D context. Extensive experiments on our dataset demonstrate the effectiveness of Visual Acoustic Fields in generating plausible impact sounds and accurately localizing impact sources. Our project page is at https://yuelei0428.github.io/projects/Visual-Acoustic-Fields/.
format Preprint
id arxiv_https___arxiv_org_abs_2503_24270
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Visual Acoustic Fields
Li, Yuelei
Kim, Hyunjin
Zhan, Fangneng
Qiu, Ri-Zhao
Ji, Mazeyu
Shan, Xiaojun
Zou, Xueyan
Liang, Paul
Pfister, Hanspeter
Wang, Xiaolong
Computer Vision and Pattern Recognition
Artificial Intelligence
Objects produce different sounds when hit, and humans can intuitively infer how an object might sound based on its appearance and material properties. Inspired by this intuition, we propose Visual Acoustic Fields, a framework that bridges hitting sounds and visual signals within a 3D space using 3D Gaussian Splatting (3DGS). Our approach features two key modules: sound generation and sound localization. The sound generation module leverages a conditional diffusion model, which takes multiscale features rendered from a feature-augmented 3DGS to generate realistic hitting sounds. Meanwhile, the sound localization module enables querying the 3D scene, represented by the feature-augmented 3DGS, to localize hitting positions based on the sound sources. To support this framework, we introduce a novel pipeline for collecting scene-level visual-sound sample pairs, achieving alignment between captured images, impact locations, and corresponding sounds. To the best of our knowledge, this is the first dataset to connect visual and acoustic signals in a 3D context. Extensive experiments on our dataset demonstrate the effectiveness of Visual Acoustic Fields in generating plausible impact sounds and accurately localizing impact sources. Our project page is at https://yuelei0428.github.io/projects/Visual-Acoustic-Fields/.
title Visual Acoustic Fields
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2503.24270