GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911562901487616 |
|---|---|
| author | Rong, Xiaobin Wang, Yushi Wang, Zheng Lu, Jing |
| author_facet | Rong, Xiaobin Wang, Yushi Wang, Zheng Lu, Jing |
| contents | We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge. The system integrates a generative branch, which performs full-stack speech restoration in a self-supervised representation domain and reconstructs the waveform via a neural vocoder, along with a predictive branch that performs spectrogram-domain enhancement, providing complementary cues. Outputs from both branches are fused by a post-processing module, which also performs bandwidth extension to generate the enhanced waveform at 48 kHz, later downsampled to the original sampling rate. This generative-predictive fusion improves robustness and perceptual quality, achieving top performance in the blind-test phase and ranking 1st in the objective evaluation. Audio examples are available at https://xiaobin-rong.github.io/gap-urgenet_demo. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_01832 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement Rong, Xiaobin Wang, Yushi Wang, Zheng Lu, Jing Audio and Speech Processing Sound We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge. The system integrates a generative branch, which performs full-stack speech restoration in a self-supervised representation domain and reconstructs the waveform via a neural vocoder, along with a predictive branch that performs spectrogram-domain enhancement, providing complementary cues. Outputs from both branches are fused by a post-processing module, which also performs bandwidth extension to generate the enhanced waveform at 48 kHz, later downsampled to the original sampling rate. This generative-predictive fusion improves robustness and perceptual quality, achieving top performance in the blind-test phase and ranking 1st in the objective evaluation. Audio examples are available at https://xiaobin-rong.github.io/gap-urgenet_demo. |
| title | GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2604.01832 |