UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Gan, Chong-Xin, Bell, Peter, Mak, Man-Wai, Li, Zhe, Jin, Zezhong, Huang, Zilong, Lee, Kong Aik
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917444049698816
author Gan, Chong-Xin
Bell, Peter
Mak, Man-Wai
Li, Zhe
Jin, Zezhong
Huang, Zilong
Lee, Kong Aik
author_facet Gan, Chong-Xin
Bell, Peter
Mak, Man-Wai
Li, Zhe
Jin, Zezhong
Huang, Zilong
Lee, Kong Aik
contents The joint training of speech enhancement and speaker embedding networks for speaker recognition is widely adopted under noisy acoustic environments. While effective, this paradigm often fails to leverage the generalization and robustness benefits inherent in large-scale speech enhancement pre-training. Moreover, maintaining the speaker information in the denoised speech is not an explicit objective of the speech enhancement process. To address these limitations, we proposed a scalable \textbf{U}Net-based \textbf{F}usion framework (UF-EMA) that considers the noisy and enhanced speech as a multi-channel input, thereby enabling the speaker encoder to exploit speaker information effectively. In addition, an \textbf{E}xponential \textbf{M}oving \textbf{A}verage strategy is applied to a speaker encoder pre-trained on clean speech to mitigate overfitting and facilitate a smooth transition from clean to noisy conditions. Experimental results on multiple noise-contaminated test sets showcase the superiority of the proposed approach.
format Preprint
id arxiv_https___arxiv_org_abs_2604_25624
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition
Gan, Chong-Xin
Bell, Peter
Mak, Man-Wai
Li, Zhe
Jin, Zezhong
Huang, Zilong
Lee, Kong Aik
Audio and Speech Processing
The joint training of speech enhancement and speaker embedding networks for speaker recognition is widely adopted under noisy acoustic environments. While effective, this paradigm often fails to leverage the generalization and robustness benefits inherent in large-scale speech enhancement pre-training. Moreover, maintaining the speaker information in the denoised speech is not an explicit objective of the speech enhancement process. To address these limitations, we proposed a scalable \textbf{U}Net-based \textbf{F}usion framework (UF-EMA) that considers the noisy and enhanced speech as a multi-channel input, thereby enabling the speaker encoder to exploit speaker information effectively. In addition, an \textbf{E}xponential \textbf{M}oving \textbf{A}verage strategy is applied to a speaker encoder pre-trained on clean speech to mitigate overfitting and facilitate a smooth transition from clean to noisy conditions. Experimental results on multiple noise-contaminated test sets showcase the superiority of the proposed approach.
title UNet-Based Fusion and Exponential Moving Average Adaptation for Noise-Robust Speaker Recognition
topic Audio and Speech Processing
url https://arxiv.org/abs/2604.25624