High-Fidelity Music Vocoder using Neural Audio Codecs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lanzendörfer, Luca A., Grötschla, Florian, Ungersböck, Michael, Wattenhofer, Roger
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917927463157760
author Lanzendörfer, Luca A.
Grötschla, Florian
Ungersböck, Michael
Wattenhofer, Roger
author_facet Lanzendörfer, Luca A.
Grötschla, Florian
Ungersböck, Michael
Wattenhofer, Roger
contents While neural vocoders have made significant progress in high-fidelity speech synthesis, their application on polyphonic music has remained underexplored. In this work, we propose DisCoder, a neural vocoder that leverages a generative adversarial encoder-decoder architecture informed by a neural audio codec to reconstruct high-fidelity 44.1 kHz audio from mel spectrograms. Our approach first transforms the mel spectrogram into a lower-dimensional representation aligned with the Descript Audio Codec (DAC) latent space before reconstructing it to an audio signal using a fine-tuned DAC decoder. DisCoder achieves state-of-the-art performance in music synthesis on several objective metrics and in a MUSHRA listening study. Our approach also shows competitive performance in speech synthesis, highlighting its potential as a universal vocoder.
format Preprint
id arxiv_https___arxiv_org_abs_2502_12759
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle High-Fidelity Music Vocoder using Neural Audio Codecs
Lanzendörfer, Luca A.
Grötschla, Florian
Ungersböck, Michael
Wattenhofer, Roger
Sound
Machine Learning
While neural vocoders have made significant progress in high-fidelity speech synthesis, their application on polyphonic music has remained underexplored. In this work, we propose DisCoder, a neural vocoder that leverages a generative adversarial encoder-decoder architecture informed by a neural audio codec to reconstruct high-fidelity 44.1 kHz audio from mel spectrograms. Our approach first transforms the mel spectrogram into a lower-dimensional representation aligned with the Descript Audio Codec (DAC) latent space before reconstructing it to an audio signal using a fine-tuned DAC decoder. DisCoder achieves state-of-the-art performance in music synthesis on several objective metrics and in a MUSHRA listening study. Our approach also shows competitive performance in speech synthesis, highlighting its potential as a universal vocoder.
title High-Fidelity Music Vocoder using Neural Audio Codecs
topic Sound
Machine Learning
url https://arxiv.org/abs/2502.12759