GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rong, Xiaobin, Wang, Yushi, Wang, Zheng, Lu, Jing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911562901487616
author Rong, Xiaobin
Wang, Yushi
Wang, Zheng
Lu, Jing
author_facet Rong, Xiaobin
Wang, Yushi
Wang, Zheng
Lu, Jing
contents We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge. The system integrates a generative branch, which performs full-stack speech restoration in a self-supervised representation domain and reconstructs the waveform via a neural vocoder, along with a predictive branch that performs spectrogram-domain enhancement, providing complementary cues. Outputs from both branches are fused by a post-processing module, which also performs bandwidth extension to generate the enhanced waveform at 48 kHz, later downsampled to the original sampling rate. This generative-predictive fusion improves robustness and perceptual quality, achieving top performance in the blind-test phase and ranking 1st in the objective evaluation. Audio examples are available at https://xiaobin-rong.github.io/gap-urgenet_demo.
format Preprint
id arxiv_https___arxiv_org_abs_2604_01832
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement
Rong, Xiaobin
Wang, Yushi
Wang, Zheng
Lu, Jing
Audio and Speech Processing
Sound
We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge. The system integrates a generative branch, which performs full-stack speech restoration in a self-supervised representation domain and reconstructs the waveform via a neural vocoder, along with a predictive branch that performs spectrogram-domain enhancement, providing complementary cues. Outputs from both branches are fused by a post-processing module, which also performs bandwidth extension to generate the enhanced waveform at 48 kHz, later downsampled to the original sampling rate. This generative-predictive fusion improves robustness and perceptual quality, achieving top performance in the blind-test phase and ranking 1st in the objective evaluation. Audio examples are available at https://xiaobin-rong.github.io/gap-urgenet_demo.
title GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2604.01832