CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Yuanhong, Shimada, Kazuki, Simon, Christian, Ikemiya, Yukara, Shibuya, Takashi, Mitsufuji, Yuki
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908479955927040
author Chen, Yuanhong
Shimada, Kazuki
Simon, Christian
Ikemiya, Yukara
Shibuya, Takashi
Mitsufuji, Yuki
author_facet Chen, Yuanhong
Shimada, Kazuki
Simon, Christian
Ikemiya, Yukara
Shibuya, Takashi
Mitsufuji, Yuki
contents Binaural audio generation (BAG) aims to convert monaural audio to stereo audio using visual prompts, requiring a deep understanding of spatial and semantic information. However, current models risk overfitting to room environments and lose fine-grained spatial details. In this paper, we propose a new audio-visual binaural generation model incorporating an audio-visual conditional normalisation layer that dynamically aligns the mean and variance of the target difference audio features using visual context, along with a new contrastive learning method to enhance spatial sensitivity by mining negative samples from shuffled visual features. We also introduce a cost-efficient way to utilise test-time augmentation in video data to enhance performance. Our approach achieves state-of-the-art generation accuracy on the FAIR-Play and MUSIC-Stereo benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02786
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
Chen, Yuanhong
Shimada, Kazuki
Simon, Christian
Ikemiya, Yukara
Shibuya, Takashi
Mitsufuji, Yuki
Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
Binaural audio generation (BAG) aims to convert monaural audio to stereo audio using visual prompts, requiring a deep understanding of spatial and semantic information. However, current models risk overfitting to room environments and lose fine-grained spatial details. In this paper, we propose a new audio-visual binaural generation model incorporating an audio-visual conditional normalisation layer that dynamically aligns the mean and variance of the target difference audio features using visual context, along with a new contrastive learning method to enhance spatial sensitivity by mining negative samples from shuffled visual features. We also introduce a cost-efficient way to utilise test-time augmentation in video data to enhance performance. Our approach achieves state-of-the-art generation accuracy on the FAIR-Play and MUSIC-Stereo benchmarks.
title CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
topic Sound
Computer Vision and Pattern Recognition
Audio and Speech Processing
url https://arxiv.org/abs/2501.02786