MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Yike, Kang, Boyi, Wang, Ziqian, Li, Xingchen, Zhang, Zihan, Li, Wenjie, Xiao, Longshuai, Xue, Wei, Xie, Lei
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914066574868480
author Zhu, Yike
Kang, Boyi
Wang, Ziqian
Li, Xingchen
Zhang, Zihan
Li, Wenjie
Xiao, Longshuai
Xue, Wei
Xie, Lei
author_facet Zhu, Yike
Kang, Boyi
Wang, Ziqian
Li, Xingchen
Zhang, Zihan
Li, Wenjie
Xiao, Longshuai
Xue, Wei
Xie, Lei
contents Speech enhancement (SE) recovers clean speech from noisy signals and is vital for applications such as telecommunications and automatic speech recognition (ASR). While generative approaches achieve strong perceptual quality, they often rely on multi-step sampling (diffusion/flow-matching) or large language models, limiting real-time deployment. To mitigate these constraints, we present MeanFlowSE, a one-step generative SE framework. It adopts MeanFlow to predict an average-velocity field for one-step latent refinement and conditions the model on self-supervised learning (SSL) representations rather than VAE latents. This design accelerates inference and provides robust acoustic-semantic guidance during training. In the Interspeech 2020 DNS Challenge blind test set and simulated test set, MeanFlowSE attains state-of-the-art (SOTA) level perceptual quality and competitive intelligibility while significantly lowering both real-time factor (RTF) and model size compared with recent generative competitors, making it suitable for practical use. The code will be released upon publication at https://github.com/Hello3orld/MeanFlowSE.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23299
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
Zhu, Yike
Kang, Boyi
Wang, Ziqian
Li, Xingchen
Zhang, Zihan
Li, Wenjie
Xiao, Longshuai
Xue, Wei
Xie, Lei
Sound
Audio and Speech Processing
Speech enhancement (SE) recovers clean speech from noisy signals and is vital for applications such as telecommunications and automatic speech recognition (ASR). While generative approaches achieve strong perceptual quality, they often rely on multi-step sampling (diffusion/flow-matching) or large language models, limiting real-time deployment. To mitigate these constraints, we present MeanFlowSE, a one-step generative SE framework. It adopts MeanFlow to predict an average-velocity field for one-step latent refinement and conditions the model on self-supervised learning (SSL) representations rather than VAE latents. This design accelerates inference and provides robust acoustic-semantic guidance during training. In the Interspeech 2020 DNS Challenge blind test set and simulated test set, MeanFlowSE attains state-of-the-art (SOTA) level perceptual quality and competitive intelligibility while significantly lowering both real-time factor (RTF) and model size compared with recent generative competitors, making it suitable for practical use. The code will be released upon publication at https://github.com/Hello3orld/MeanFlowSE.
title MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2509.23299