A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Ningyuan, Li, Yize, Cuji, Diego A., Corey, Ryan M., Zhao, Pu, Lin, Xue, Singer, Andrew C.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909054844010496
author Yang, Ningyuan
Li, Yize
Cuji, Diego A.
Corey, Ryan M.
Zhao, Pu
Lin, Xue
Singer, Andrew C.
author_facet Yang, Ningyuan
Li, Yize
Cuji, Diego A.
Corey, Ryan M.
Zhao, Pu
Lin, Xue
Singer, Andrew C.
contents Audio super-resolution (SR), also referred to as bandwidth extension (BWE), aims to reconstruct high-fidelity signals from low-resolution (LR) or band-limited (BL) observations, an inherently ill-posed task due to the ambiguity of missing high-frequency (HF) content. This survey provides a comprehensive overview of the field, with a particular focus on the paradigm shift from discriminative mapping to modern generative modeling. We first review early discriminative deep neural network (DNN) models, which formulate BWE/SR as a deterministic mapping problem and are prone to regression-to-the-mean effects and spectral over-smoothing. We then systematically review generative approaches, including autoregressive (AR) models, variational autoencoders (VAEs), generative adversarial networks (GANs), diffusion and score-based models, flow-based methods, and Schrödinger bridges. Across these approaches, we examine key design aspects, including representation domain, architecture, conditioning mechanisms, and trade-offs among reconstruction fidelity, perceptual quality, robustness, and computational efficiency. Furthermore, we discuss emerging directions involving large language models (LLMs) and multimodal foundation models, and highlight open challenges in perceptual evaluation, phase modeling, and real-world generalization. By providing a structured taxonomy and unified perspective, this survey establishes a comprehensive foundation and offers a practical roadmap for advancing BWE/SR from deterministic point estimation toward distribution-aware generative modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16681
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models
Yang, Ningyuan
Li, Yize
Cuji, Diego A.
Corey, Ryan M.
Zhao, Pu
Lin, Xue
Singer, Andrew C.
Audio and Speech Processing
Sound
Signal Processing
Audio super-resolution (SR), also referred to as bandwidth extension (BWE), aims to reconstruct high-fidelity signals from low-resolution (LR) or band-limited (BL) observations, an inherently ill-posed task due to the ambiguity of missing high-frequency (HF) content. This survey provides a comprehensive overview of the field, with a particular focus on the paradigm shift from discriminative mapping to modern generative modeling. We first review early discriminative deep neural network (DNN) models, which formulate BWE/SR as a deterministic mapping problem and are prone to regression-to-the-mean effects and spectral over-smoothing. We then systematically review generative approaches, including autoregressive (AR) models, variational autoencoders (VAEs), generative adversarial networks (GANs), diffusion and score-based models, flow-based methods, and Schrödinger bridges. Across these approaches, we examine key design aspects, including representation domain, architecture, conditioning mechanisms, and trade-offs among reconstruction fidelity, perceptual quality, robustness, and computational efficiency. Furthermore, we discuss emerging directions involving large language models (LLMs) and multimodal foundation models, and highlight open challenges in perceptual evaluation, phase modeling, and real-world generalization. By providing a structured taxonomy and unified perspective, this survey establishes a comprehensive foundation and offers a practical roadmap for advancing BWE/SR from deterministic point estimation toward distribution-aware generative modeling.
title A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models
topic Audio and Speech Processing
Sound
Signal Processing
url https://arxiv.org/abs/2605.16681