A Hybrid Discriminative and Generative System for Universal Speech Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yinghao, Liu, Chengwei, Liang, Xiaotao, Yan, Haoyin, Xue, Shaofei, Xue, Zheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912851757629440
author Liu, Yinghao
Liu, Chengwei
Liang, Xiaotao
Yan, Haoyin
Xue, Shaofei
Xue, Zheng
author_facet Liu, Yinghao
Liu, Chengwei
Liang, Xiaotao
Yan, Haoyin
Xue, Shaofei
Xue, Zheng
contents Universal speech enhancement aims at handling inputs with various speech distortions and recording conditions. In this work, we propose a novel hybrid architecture that synergizes the signal fidelity of discriminative modeling with the reconstruction capabilities of generative modeling. Our system utilizes the discriminative TF-GridNet model with the Sampling-Frequency-Independent strategy to handle variable sampling rates universally. In parallel, an autoregressive model combined with spectral mapping modeling generates detail-rich speech while effectively suppressing generative artifacts. Finally, a fusion network learns adaptive weights of the two outputs under the optimization of signal-level losses and the comprehensive Speech Quality Assessment (SQA) loss. Our proposed system is evaluated in the ICASSP 2026 URGENT Challenge (Track 1) and ranks the third place.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19113
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Hybrid Discriminative and Generative System for Universal Speech Enhancement
Liu, Yinghao
Liu, Chengwei
Liang, Xiaotao
Yan, Haoyin
Xue, Shaofei
Xue, Zheng
Sound
Audio and Speech Processing
Universal speech enhancement aims at handling inputs with various speech distortions and recording conditions. In this work, we propose a novel hybrid architecture that synergizes the signal fidelity of discriminative modeling with the reconstruction capabilities of generative modeling. Our system utilizes the discriminative TF-GridNet model with the Sampling-Frequency-Independent strategy to handle variable sampling rates universally. In parallel, an autoregressive model combined with spectral mapping modeling generates detail-rich speech while effectively suppressing generative artifacts. Finally, a fusion network learns adaptive weights of the two outputs under the optimization of signal-level losses and the comprehensive Speech Quality Assessment (SQA) loss. Our proposed system is evaluated in the ICASSP 2026 URGENT Challenge (Track 1) and ranks the third place.
title A Hybrid Discriminative and Generative System for Universal Speech Enhancement
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2601.19113