SAGA-SR: Semantically and Acoustically Guided Audio Super-Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Im, Jaekwon, Nam, Juhan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909814788980736
author Im, Jaekwon
Nam, Juhan
author_facet Im, Jaekwon
Nam, Juhan
contents Versatile audio super-resolution (SR) aims to predict high-frequency components from low-resolution audio across diverse domains such as speech, music, and sound effects. Existing diffusion-based SR methods often fail to produce semantically aligned outputs and struggle with consistent high-frequency reconstruction. In this paper, we propose SAGA-SR, a versatile audio SR model that combines semantic and acoustic guidance. Based on a DiT backbone trained with a flow matching objective, SAGA-SR is conditioned on text and spectral roll-off embeddings. Due to the effective guidance provided by its conditioning, SAGA-SR robustly upsamples audio from arbitrary input sampling rates between 4 kHz and 32 kHz to 44.1 kHz. Both objective and subjective evaluations show that SAGA-SR achieves state-of-the-art performance across all test cases. Sound examples and code for the proposed model are available online.
format Preprint
id arxiv_https___arxiv_org_abs_2509_24924
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SAGA-SR: Semantically and Acoustically Guided Audio Super-Resolution
Im, Jaekwon
Nam, Juhan
Audio and Speech Processing
Versatile audio super-resolution (SR) aims to predict high-frequency components from low-resolution audio across diverse domains such as speech, music, and sound effects. Existing diffusion-based SR methods often fail to produce semantically aligned outputs and struggle with consistent high-frequency reconstruction. In this paper, we propose SAGA-SR, a versatile audio SR model that combines semantic and acoustic guidance. Based on a DiT backbone trained with a flow matching objective, SAGA-SR is conditioned on text and spectral roll-off embeddings. Due to the effective guidance provided by its conditioning, SAGA-SR robustly upsamples audio from arbitrary input sampling rates between 4 kHz and 32 kHz to 44.1 kHz. Both objective and subjective evaluations show that SAGA-SR achieves state-of-the-art performance across all test cases. Sound examples and code for the proposed model are available online.
title SAGA-SR: Semantically and Acoustically Guided Audio Super-Resolution
topic Audio and Speech Processing
url https://arxiv.org/abs/2509.24924