FlowSE: Flow Matching-based Speech Enhancement

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Lee, Seonggyu, Cheong, Sein, Han, Sangwook, Shin, Jong Won
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913982687739904
author Lee, Seonggyu
Cheong, Sein
Han, Sangwook
Shin, Jong Won
author_facet Lee, Seonggyu
Cheong, Sein
Han, Sangwook
Shin, Jong Won
contents Diffusion probabilistic models have shown impressive performance for speech enhancement, but they typically require 25 to 60 function evaluations in the inference phase, resulting in heavy computational complexity. Recently, a fine-tuning method was proposed to correct the reverse process, which significantly lowered the number of function evaluations (NFE). Flow matching is a method to train continuous normalizing flows which model probability paths from known distributions to unknown distributions including those described by diffusion processes. In this paper, we propose a speech enhancement based on conditional flow matching. The proposed method achieved the performance comparable to those for the diffusion-based speech enhancement with the NFE of 60 when the NFE was 5, and showed similar performance with the diffusion model correcting the reverse process at the same NFE from 1 to 5 without additional fine tuning procedure. We also have shown that the corresponding diffusion model derived from the conditional probability path with a modified optimal transport conditional vector field demonstrated similar performances with the NFE of 5 without any fine-tuning procedure.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06840
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FlowSE: Flow Matching-based Speech Enhancement
Lee, Seonggyu
Cheong, Sein
Han, Sangwook
Shin, Jong Won
Audio and Speech Processing
Signal Processing
Diffusion probabilistic models have shown impressive performance for speech enhancement, but they typically require 25 to 60 function evaluations in the inference phase, resulting in heavy computational complexity. Recently, a fine-tuning method was proposed to correct the reverse process, which significantly lowered the number of function evaluations (NFE). Flow matching is a method to train continuous normalizing flows which model probability paths from known distributions to unknown distributions including those described by diffusion processes. In this paper, we propose a speech enhancement based on conditional flow matching. The proposed method achieved the performance comparable to those for the diffusion-based speech enhancement with the NFE of 60 when the NFE was 5, and showed similar performance with the diffusion model correcting the reverse process at the same NFE from 1 to 5 without additional fine tuning procedure. We also have shown that the corresponding diffusion model derived from the conditional probability path with a modified optimal transport conditional vector field demonstrated similar performances with the NFE of 5 without any fine-tuning procedure.
title FlowSE: Flow Matching-based Speech Enhancement
topic Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2508.06840