Mutual Learning for Acoustic Matching and Dereverberation via Visual Scene-driven Diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Jian, Wang, Wenguan, Yang, Yi, Zheng, Feng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913431329701888
author Ma, Jian
Wang, Wenguan
Yang, Yi
Zheng, Feng
author_facet Ma, Jian
Wang, Wenguan
Yang, Yi
Zheng, Feng
contents Visual acoustic matching (VAM) is pivotal for enhancing the immersive experience, and the task of dereverberation is effective in improving audio intelligibility. Existing methods treat each task independently, overlooking the inherent reciprocity between them. Moreover, these methods depend on paired training data, which is challenging to acquire, impeding the utilization of extensive unpaired data. In this paper, we introduce MVSD, a mutual learning framework based on diffusion models. MVSD considers the two tasks symmetrically, exploiting the reciprocal relationship to facilitate learning from inverse tasks and overcome data scarcity. Furthermore, we employ the diffusion model as foundational conditional converters to circumvent the training instability and over-smoothing drawbacks of conventional GAN architectures. Specifically, MVSD employs two converters: one for VAM called reverberator and one for dereverberation called dereverberator. The dereverberator judges whether the reverberation audio generated by reverberator sounds like being in the conditional visual scenario, and vice versa. By forming a closed loop, these two converters can generate informative feedback signals to optimize the inverse tasks, even with easily acquired one-way unpaired data. Extensive experiments on two standard benchmarks, i.e., SoundSpaces-Speech and Acoustic AVSpeech, exhibit that our framework can improve the performance of the reverberator and dereverberator and better match specified visual scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10373
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mutual Learning for Acoustic Matching and Dereverberation via Visual Scene-driven Diffusion
Ma, Jian
Wang, Wenguan
Yang, Yi
Zheng, Feng
Sound
Artificial Intelligence
Computer Vision and Pattern Recognition
Audio and Speech Processing
Visual acoustic matching (VAM) is pivotal for enhancing the immersive experience, and the task of dereverberation is effective in improving audio intelligibility. Existing methods treat each task independently, overlooking the inherent reciprocity between them. Moreover, these methods depend on paired training data, which is challenging to acquire, impeding the utilization of extensive unpaired data. In this paper, we introduce MVSD, a mutual learning framework based on diffusion models. MVSD considers the two tasks symmetrically, exploiting the reciprocal relationship to facilitate learning from inverse tasks and overcome data scarcity. Furthermore, we employ the diffusion model as foundational conditional converters to circumvent the training instability and over-smoothing drawbacks of conventional GAN architectures. Specifically, MVSD employs two converters: one for VAM called reverberator and one for dereverberation called dereverberator. The dereverberator judges whether the reverberation audio generated by reverberator sounds like being in the conditional visual scenario, and vice versa. By forming a closed loop, these two converters can generate informative feedback signals to optimize the inverse tasks, even with easily acquired one-way unpaired data. Extensive experiments on two standard benchmarks, i.e., SoundSpaces-Speech and Acoustic AVSpeech, exhibit that our framework can improve the performance of the reverberator and dereverberator and better match specified visual scenarios.
title Mutual Learning for Acoustic Matching and Dereverberation via Visual Scene-driven Diffusion
topic Sound
Artificial Intelligence
Computer Vision and Pattern Recognition
Audio and Speech Processing
url https://arxiv.org/abs/2407.10373