Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Torres, Bernardo, Moussallam, Manuel, Meseguer-Brocal, Gabriel
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908792938037248
author Torres, Bernardo
Moussallam, Manuel
Meseguer-Brocal, Gabriel
author_facet Torres, Bernardo
Moussallam, Manuel
Meseguer-Brocal, Gabriel
contents Audio autoencoders learn useful, compressed audio representations, but their non-linear latent spaces prevent intuitive algebraic manipulation such as mixing or scaling. We introduce a simple training methodology to induce linearity in a high-compression Consistency Autoencoder (CAE) by using data augmentation, thereby inducing homogeneity (equivariance to scalar gain) and additivity (the decoder preserves addition) without altering the model's architecture or loss function. When trained with our method, the CAE exhibits linear behavior in both the encoder and decoder while preserving reconstruction fidelity. We test the practical utility of our learned space on music source composition and separation via simple latent arithmetic. This work presents a straightforward technique for constructing structured latent spaces, enabling more intuitive and efficient audio processing.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23530
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization
Torres, Bernardo
Moussallam, Manuel
Meseguer-Brocal, Gabriel
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Audio autoencoders learn useful, compressed audio representations, but their non-linear latent spaces prevent intuitive algebraic manipulation such as mixing or scaling. We introduce a simple training methodology to induce linearity in a high-compression Consistency Autoencoder (CAE) by using data augmentation, thereby inducing homogeneity (equivariance to scalar gain) and additivity (the decoder preserves addition) without altering the model's architecture or loss function. When trained with our method, the CAE exhibits linear behavior in both the encoder and decoder while preserving reconstruction fidelity. We test the practical utility of our learned space on music source composition and separation via simple latent arithmetic. This work presents a straightforward technique for constructing structured latent spaces, enabling more intuitive and efficient audio processing.
title Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2510.23530