In-the-wild Audio Spatialization with Flexible Text-guided Localization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pan, Tianrui, Liu, Jie, Huang, Zewen, Tang, Jie, Wu, Gangshan
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913869936459776
author Pan, Tianrui
Liu, Jie
Huang, Zewen
Tang, Jie
Wu, Gangshan
author_facet Pan, Tianrui
Liu, Jie
Huang, Zewen
Tang, Jie
Wu, Gangshan
contents To enhance immersive experiences, binaural audio offers spatial awareness of sounding objects in AR, VR, and embodied AI applications. While existing audio spatialization methods can generally map any available monaural audio to binaural audio signals, they often lack the flexible and interactive control needed in complex multi-object user-interactive environments. To address this, we propose a Text-guided Audio Spatialization (TAS) framework that utilizes flexible text prompts and evaluates our model from unified generation and comprehension perspectives. Due to the limited availability of premium and large-scale stereo data, we construct the SpatialTAS dataset, which encompasses 376,000 simulated binaural audio samples to facilitate the training of our model. Our model learns binaural differences guided by 3D spatial location and relative position prompts, augmented by flipped-channel audio. It outperforms existing methods on both simulated and real-recorded datasets, demonstrating superior generalization and accuracy. Besides, we develop an assessment model based on Llama-3.1-8B, which evaluates the spatial semantic coherence between our generated binaural audio and text prompts through a spatial reasoning task. Results demonstrate that text prompts provide flexible and interactive control to generate binaural audio with excellent quality and semantic consistency in spatial locations. Dataset is available at \href{https://github.com/Alice01010101/TASU}
format Preprint
id arxiv_https___arxiv_org_abs_2506_00927
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle In-the-wild Audio Spatialization with Flexible Text-guided Localization
Pan, Tianrui
Liu, Jie
Huang, Zewen
Tang, Jie
Wu, Gangshan
Sound
Artificial Intelligence
Audio and Speech Processing
To enhance immersive experiences, binaural audio offers spatial awareness of sounding objects in AR, VR, and embodied AI applications. While existing audio spatialization methods can generally map any available monaural audio to binaural audio signals, they often lack the flexible and interactive control needed in complex multi-object user-interactive environments. To address this, we propose a Text-guided Audio Spatialization (TAS) framework that utilizes flexible text prompts and evaluates our model from unified generation and comprehension perspectives. Due to the limited availability of premium and large-scale stereo data, we construct the SpatialTAS dataset, which encompasses 376,000 simulated binaural audio samples to facilitate the training of our model. Our model learns binaural differences guided by 3D spatial location and relative position prompts, augmented by flipped-channel audio. It outperforms existing methods on both simulated and real-recorded datasets, demonstrating superior generalization and accuracy. Besides, we develop an assessment model based on Llama-3.1-8B, which evaluates the spatial semantic coherence between our generated binaural audio and text prompts through a spatial reasoning task. Results demonstrate that text prompts provide flexible and interactive control to generate binaural audio with excellent quality and semantic consistency in spatial locations. Dataset is available at \href{https://github.com/Alice01010101/TASU}
title In-the-wild Audio Spatialization with Flexible Text-guided Localization
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2506.00927