EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cerovaz, Luca, Mancusi, Michele, Rodolà, Emanuele
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908793028214784
author Cerovaz, Luca
Mancusi, Michele
Rodolà, Emanuele
author_facet Cerovaz, Luca
Mancusi, Michele
Rodolà, Emanuele
contents Audio codecs power discrete music generative modelling, music streaming and immersive media by shrinking PCM audio to bandwidth-friendly bit-rates. Recent works have gravitated towards processing in the spectral domain; however, spectrogram-domains typically struggle with phase modeling which is naturally complex-valued. Most frequency-domain neural codecs either disregard phase information or encode it as two separate real-valued channels, limiting spatial fidelity. This entails the need to introduce adversarial discriminators at the expense of convergence speed and training stability to compensate for the inadequate representation power of the audio signal. In this work we introduce an end-to-end complex-valued RVQ-VAE audio codec that preserves magnitude-phase coupling across the entire analysis-quantization-synthesis pipeline and removes adversarial discriminators and diffusion post-filters. Without GANs or diffusion we match or surpass much longer-trained baselines in-domain and reach SOTA out-of-domain performance. Compared to standard baselines that train for hundreds of thousands of steps, our model reducing training budget by an order of magnitude is markedly more compute-efficient while preserving high perceptual quality.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17517
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding
Cerovaz, Luca
Mancusi, Michele
Rodolà, Emanuele
Sound
Machine Learning
Audio and Speech Processing
Audio codecs power discrete music generative modelling, music streaming and immersive media by shrinking PCM audio to bandwidth-friendly bit-rates. Recent works have gravitated towards processing in the spectral domain; however, spectrogram-domains typically struggle with phase modeling which is naturally complex-valued. Most frequency-domain neural codecs either disregard phase information or encode it as two separate real-valued channels, limiting spatial fidelity. This entails the need to introduce adversarial discriminators at the expense of convergence speed and training stability to compensate for the inadequate representation power of the audio signal. In this work we introduce an end-to-end complex-valued RVQ-VAE audio codec that preserves magnitude-phase coupling across the entire analysis-quantization-synthesis pipeline and removes adversarial discriminators and diffusion post-filters. Without GANs or diffusion we match or surpass much longer-trained baselines in-domain and reach SOTA out-of-domain performance. Compared to standard baselines that train for hundreds of thousands of steps, our model reducing training budget by an order of magnitude is markedly more compute-efficient while preserving high perceptual quality.
title EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2601.17517