Learning to Upsample and Upmix Audio in the Latent Domain

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bralios, Dimitrios, Smaragdis, Paris, Casebeer, Jonah
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909776398516224
author Bralios, Dimitrios
Smaragdis, Paris
Casebeer, Jonah
author_facet Bralios, Dimitrios
Smaragdis, Paris
Casebeer, Jonah
contents Neural audio autoencoders create compact latent representations that preserve perceptually important information, serving as the foundation for both modern audio compression systems and generation approaches like next-token prediction and latent diffusion. Despite their prevalence, most audio processing operations, such as spatial and spectral up-sampling, still inefficiently operate on raw waveforms or spectral representations rather than directly on these compressed representations. We propose a framework that performs audio processing operations entirely within an autoencoder's latent space, eliminating the need to decode to raw audio formats. Our approach dramatically simplifies training by operating solely in the latent domain, with a latent L1 reconstruction term, augmented by a single latent adversarial discriminator. This contrasts sharply with raw-audio methods that typically require complex combinations of multi-scale losses and discriminators. Through experiments in bandwidth extension and mono-to-stereo up-mixing, we demonstrate computational efficiency gains of up to 100x while maintaining quality comparable to post-processing on raw audio. This work establishes a more efficient paradigm for audio processing pipelines that already incorporate autoencoders, enabling significantly faster and more resource-efficient workflows across various audio tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00681
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning to Upsample and Upmix Audio in the Latent Domain
Bralios, Dimitrios
Smaragdis, Paris
Casebeer, Jonah
Sound
Machine Learning
Audio and Speech Processing
Neural audio autoencoders create compact latent representations that preserve perceptually important information, serving as the foundation for both modern audio compression systems and generation approaches like next-token prediction and latent diffusion. Despite their prevalence, most audio processing operations, such as spatial and spectral up-sampling, still inefficiently operate on raw waveforms or spectral representations rather than directly on these compressed representations. We propose a framework that performs audio processing operations entirely within an autoencoder's latent space, eliminating the need to decode to raw audio formats. Our approach dramatically simplifies training by operating solely in the latent domain, with a latent L1 reconstruction term, augmented by a single latent adversarial discriminator. This contrasts sharply with raw-audio methods that typically require complex combinations of multi-scale losses and discriminators. Through experiments in bandwidth extension and mono-to-stereo up-mixing, we demonstrate computational efficiency gains of up to 100x while maintaining quality comparable to post-processing on raw audio. This work establishes a more efficient paradigm for audio processing pipelines that already incorporate autoencoders, enabling significantly faster and more resource-efficient workflows across various audio tasks.
title Learning to Upsample and Upmix Audio in the Latent Domain
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2506.00681