BUDDy: Single-Channel Blind Unsupervised Dereverberation with Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moliner, Eloi, Lemercier, Jean-Marie, Welker, Simon, Gerkmann, Timo, Välimäki, Vesa
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916238106558464
author Moliner, Eloi
Lemercier, Jean-Marie
Welker, Simon
Gerkmann, Timo
Välimäki, Vesa
author_facet Moliner, Eloi
Lemercier, Jean-Marie
Welker, Simon
Gerkmann, Timo
Välimäki, Vesa
contents In this paper, we present an unsupervised single-channel method for joint blind dereverberation and room impulse response estimation, based on posterior sampling with diffusion models. We parameterize the reverberation operator using a filter with exponential decay for each frequency subband, and iteratively estimate the corresponding parameters as the speech utterance gets refined along the reverse diffusion trajectory. A measurement consistency criterion enforces the fidelity of the generated speech with the reverberant measurement, while an unconditional diffusion model implements a strong prior for clean speech generation. Without any knowledge of the room impulse response nor any coupled reverberant-anechoic data, we can successfully perform dereverberation in various acoustic scenarios. Our method significantly outperforms previous blind unsupervised baselines, and we demonstrate its increased robustness to unseen acoustic conditions in comparison to blind supervised methods. Audio samples and code are available online.
format Preprint
id arxiv_https___arxiv_org_abs_2405_04272
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BUDDy: Single-Channel Blind Unsupervised Dereverberation with Diffusion Models
Moliner, Eloi
Lemercier, Jean-Marie
Welker, Simon
Gerkmann, Timo
Välimäki, Vesa
Audio and Speech Processing
Machine Learning
Sound
In this paper, we present an unsupervised single-channel method for joint blind dereverberation and room impulse response estimation, based on posterior sampling with diffusion models. We parameterize the reverberation operator using a filter with exponential decay for each frequency subband, and iteratively estimate the corresponding parameters as the speech utterance gets refined along the reverse diffusion trajectory. A measurement consistency criterion enforces the fidelity of the generated speech with the reverberant measurement, while an unconditional diffusion model implements a strong prior for clean speech generation. Without any knowledge of the room impulse response nor any coupled reverberant-anechoic data, we can successfully perform dereverberation in various acoustic scenarios. Our method significantly outperforms previous blind unsupervised baselines, and we demonstrate its increased robustness to unseen acoustic conditions in comparison to blind supervised methods. Audio samples and code are available online.
title BUDDy: Single-Channel Blind Unsupervised Dereverberation with Diffusion Models
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2405.04272