CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866908479955927040 |
|---|---|
| author | Chen, Yuanhong Shimada, Kazuki Simon, Christian Ikemiya, Yukara Shibuya, Takashi Mitsufuji, Yuki |
| author_facet | Chen, Yuanhong Shimada, Kazuki Simon, Christian Ikemiya, Yukara Shibuya, Takashi Mitsufuji, Yuki |
| contents | Binaural audio generation (BAG) aims to convert monaural audio to stereo audio using visual prompts, requiring a deep understanding of spatial and semantic information. However, current models risk overfitting to room environments and lose fine-grained spatial details. In this paper, we propose a new audio-visual binaural generation model incorporating an audio-visual conditional normalisation layer that dynamically aligns the mean and variance of the target difference audio features using visual context, along with a new contrastive learning method to enhance spatial sensitivity by mining negative samples from shuffled visual features. We also introduce a cost-efficient way to utilise test-time augmentation in video data to enhance performance. Our approach achieves state-of-the-art generation accuracy on the FAIR-Play and MUSIC-Stereo benchmarks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_02786 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation Chen, Yuanhong Shimada, Kazuki Simon, Christian Ikemiya, Yukara Shibuya, Takashi Mitsufuji, Yuki Sound Computer Vision and Pattern Recognition Audio and Speech Processing Binaural audio generation (BAG) aims to convert monaural audio to stereo audio using visual prompts, requiring a deep understanding of spatial and semantic information. However, current models risk overfitting to room environments and lose fine-grained spatial details. In this paper, we propose a new audio-visual binaural generation model incorporating an audio-visual conditional normalisation layer that dynamically aligns the mean and variance of the target difference audio features using visual context, along with a new contrastive learning method to enhance spatial sensitivity by mining negative samples from shuffled visual features. We also introduce a cost-efficient way to utilise test-time augmentation in video data to enhance performance. Our approach achieves state-of-the-art generation accuracy on the FAIR-Play and MUSIC-Stereo benchmarks. |
| title | CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation |
| topic | Sound Computer Vision and Pattern Recognition Audio and Speech Processing |
| url | https://arxiv.org/abs/2501.02786 |