Diffusion Gaussian Mixture Audio Denoise

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Pu, Li, Junhui, Li, Jialu, Guo, Liangdong, Zhang, Youshan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909223276773376
author Wang, Pu
Li, Junhui
Li, Jialu
Guo, Liangdong
Zhang, Youshan
author_facet Wang, Pu
Li, Junhui
Li, Jialu
Guo, Liangdong
Zhang, Youshan
contents Recent diffusion models have achieved promising performances in audio-denoising tasks. The unique property of the reverse process could recover clean signals. However, the distribution of real-world noises does not comply with a single Gaussian distribution and is even unknown. The sampling of Gaussian noise conditions limits its application scenarios. To overcome these challenges, we propose a DiffGMM model, a denoising model based on the diffusion and Gaussian mixture models. We employ the reverse process to estimate parameters for the Gaussian mixture model. Given a noisy audio signal, we first apply a 1D-U-Net to extract features and train linear layers to estimate parameters for the Gaussian mixture model, and we approximate the real noise distributions. The noisy signal is continuously subtracted from the estimated noise to output clean audio signals. Extensive experimental results demonstrate that the proposed DiffGMM model achieves state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2406_09154
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Diffusion Gaussian Mixture Audio Denoise
Wang, Pu
Li, Junhui
Li, Jialu
Guo, Liangdong
Zhang, Youshan
Sound
Computation and Language
Audio and Speech Processing
Recent diffusion models have achieved promising performances in audio-denoising tasks. The unique property of the reverse process could recover clean signals. However, the distribution of real-world noises does not comply with a single Gaussian distribution and is even unknown. The sampling of Gaussian noise conditions limits its application scenarios. To overcome these challenges, we propose a DiffGMM model, a denoising model based on the diffusion and Gaussian mixture models. We employ the reverse process to estimate parameters for the Gaussian mixture model. Given a noisy audio signal, we first apply a 1D-U-Net to extract features and train linear layers to estimate parameters for the Gaussian mixture model, and we approximate the real noise distributions. The noisy signal is continuously subtracted from the estimated noise to output clean audio signals. Extensive experimental results demonstrate that the proposed DiffGMM model achieves state-of-the-art performance.
title Diffusion Gaussian Mixture Audio Denoise
topic Sound
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2406.09154