A Generative-First Neural Audio Autoencoder

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Casebeer, Jonah, Zhu, Ge, Wang, Zhepei, Bryan, Nicholas J.
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910027680317440
author Casebeer, Jonah
Zhu, Ge
Wang, Zhepei
Bryan, Nicholas J.
author_facet Casebeer, Jonah
Zhu, Ge
Wang, Zhepei
Bryan, Nicholas J.
contents Neural autoencoders underpin generative models. Practical, large-scale use of neural autoencoders for generative modeling necessitates fast encoding, low latent rates, and a single model across representations. Existing approaches are reconstruction-first: they incur high latent rates, slow encoding, and separate architectures for discrete vs. continuous latents and for different audio channel formats, hindering workflows from preprocessing to inference conditioning. We introduce a generative-first architecture for audio autoencoding that increases temporal downsampling from 2048x to 3360x and supports continuous and discrete representations and common audio channel formats in one model. By balancing compression, quality, and speed, it delivers 10x faster encoding, 1.6x lower rates, and eliminates channel-format-specific variants while maintaining competitive reconstruction quality. This enables applications previously constrained by processing costs: a 60-second mono signal compresses to 788 tokens, making generative modeling more tractable.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15749
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Generative-First Neural Audio Autoencoder
Casebeer, Jonah
Zhu, Ge
Wang, Zhepei
Bryan, Nicholas J.
Sound
Audio and Speech Processing
Neural autoencoders underpin generative models. Practical, large-scale use of neural autoencoders for generative modeling necessitates fast encoding, low latent rates, and a single model across representations. Existing approaches are reconstruction-first: they incur high latent rates, slow encoding, and separate architectures for discrete vs. continuous latents and for different audio channel formats, hindering workflows from preprocessing to inference conditioning. We introduce a generative-first architecture for audio autoencoding that increases temporal downsampling from 2048x to 3360x and supports continuous and discrete representations and common audio channel formats in one model. By balancing compression, quality, and speed, it delivers 10x faster encoding, 1.6x lower rates, and eliminates channel-format-specific variants while maintaining competitive reconstruction quality. This enables applications previously constrained by processing costs: a 60-second mono signal compresses to 788 tokens, making generative modeling more tractable.
title A Generative-First Neural Audio Autoencoder
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2602.15749