High-Fidelity Music Vocoder using Neural Audio Codecs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866917927463157760 |
|---|---|
| author | Lanzendörfer, Luca A. Grötschla, Florian Ungersböck, Michael Wattenhofer, Roger |
| author_facet | Lanzendörfer, Luca A. Grötschla, Florian Ungersböck, Michael Wattenhofer, Roger |
| contents | While neural vocoders have made significant progress in high-fidelity speech synthesis, their application on polyphonic music has remained underexplored. In this work, we propose DisCoder, a neural vocoder that leverages a generative adversarial encoder-decoder architecture informed by a neural audio codec to reconstruct high-fidelity 44.1 kHz audio from mel spectrograms. Our approach first transforms the mel spectrogram into a lower-dimensional representation aligned with the Descript Audio Codec (DAC) latent space before reconstructing it to an audio signal using a fine-tuned DAC decoder. DisCoder achieves state-of-the-art performance in music synthesis on several objective metrics and in a MUSHRA listening study. Our approach also shows competitive performance in speech synthesis, highlighting its potential as a universal vocoder. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_12759 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | High-Fidelity Music Vocoder using Neural Audio Codecs Lanzendörfer, Luca A. Grötschla, Florian Ungersböck, Michael Wattenhofer, Roger Sound Machine Learning While neural vocoders have made significant progress in high-fidelity speech synthesis, their application on polyphonic music has remained underexplored. In this work, we propose DisCoder, a neural vocoder that leverages a generative adversarial encoder-decoder architecture informed by a neural audio codec to reconstruct high-fidelity 44.1 kHz audio from mel spectrograms. Our approach first transforms the mel spectrogram into a lower-dimensional representation aligned with the Descript Audio Codec (DAC) latent space before reconstructing it to an audio signal using a fine-tuned DAC decoder. DisCoder achieves state-of-the-art performance in music synthesis on several objective metrics and in a MUSHRA listening study. Our approach also shows competitive performance in speech synthesis, highlighting its potential as a universal vocoder. |
| title | High-Fidelity Music Vocoder using Neural Audio Codecs |
| topic | Sound Machine Learning |
| url | https://arxiv.org/abs/2502.12759 |