Diffusion-Based Unsupervised Audio-Visual Speech Separation in Noisy Environments with Noise Prior

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yemini, Yochai, Ben-Ari, Rami, Gannot, Sharon, Fetaya, Ethan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912591392014336
author Yemini, Yochai
Ben-Ari, Rami
Gannot, Sharon
Fetaya, Ethan
author_facet Yemini, Yochai
Ben-Ari, Rami
Gannot, Sharon
Fetaya, Ethan
contents In this paper, we address the problem of single-microphone speech separation in the presence of ambient noise. We propose a generative unsupervised technique that directly models both clean speech and structured noise components, training exclusively on these individual signals rather than noisy mixtures. Our approach leverages an audio-visual score model that incorporates visual cues to serve as a strong generative speech prior. By explicitly modelling the noise distribution alongside the speech distribution, we enable effective decomposition through the inverse problem paradigm. We perform speech separation by sampling from the posterior distributions via a reverse diffusion process, which directly estimates and removes the modelled noise component to recover clean constituent signals. Experimental results demonstrate promising performance, highlighting the effectiveness of our direct noise modelling approach in challenging acoustic environments.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14379
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Diffusion-Based Unsupervised Audio-Visual Speech Separation in Noisy Environments with Noise Prior
Yemini, Yochai
Ben-Ari, Rami
Gannot, Sharon
Fetaya, Ethan
Audio and Speech Processing
Machine Learning
In this paper, we address the problem of single-microphone speech separation in the presence of ambient noise. We propose a generative unsupervised technique that directly models both clean speech and structured noise components, training exclusively on these individual signals rather than noisy mixtures. Our approach leverages an audio-visual score model that incorporates visual cues to serve as a strong generative speech prior. By explicitly modelling the noise distribution alongside the speech distribution, we enable effective decomposition through the inverse problem paradigm. We perform speech separation by sampling from the posterior distributions via a reverse diffusion process, which directly estimates and removes the modelled noise component to recover clean constituent signals. Experimental results demonstrate promising performance, highlighting the effectiveness of our direct noise modelling approach in challenging acoustic environments.
title Diffusion-Based Unsupervised Audio-Visual Speech Separation in Noisy Environments with Noise Prior
topic Audio and Speech Processing
Machine Learning
url https://arxiv.org/abs/2509.14379