Linear RNNs for autoregressive generation of long music samples

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Szewczyk, Konrad, Fernández, Daniel Gallo, Townsend, James
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915531053858816
author Szewczyk, Konrad
Fernández, Daniel Gallo
Townsend, James
author_facet Szewczyk, Konrad
Fernández, Daniel Gallo
Townsend, James
contents Directly learning to generate audio waveforms in an autoregressive manner is a challenging task, due to the length of the raw sequences and the existence of important structure on many different timescales. Traditional approaches based on recurrent neural networks, as well as causal convolutions and self-attention, have only had limited success on this task. However, recent work has shown that deep state space models, also referred to as linear RNNs, can be highly efficient in this context. In this work, we push the boundaries of linear RNNs applied to raw audio modeling, investigating the effects of different architectural choices and using context-parallelism to enable training on sequences up to one minute (1M tokens) in length. We present a model, HarmonicRNN, which attains state of the art log-likelihoods and perceptual metrics on small-scale datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02401
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Linear RNNs for autoregressive generation of long music samples
Szewczyk, Konrad
Fernández, Daniel Gallo
Townsend, James
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Directly learning to generate audio waveforms in an autoregressive manner is a challenging task, due to the length of the raw sequences and the existence of important structure on many different timescales. Traditional approaches based on recurrent neural networks, as well as causal convolutions and self-attention, have only had limited success on this task. However, recent work has shown that deep state space models, also referred to as linear RNNs, can be highly efficient in this context. In this work, we push the boundaries of linear RNNs applied to raw audio modeling, investigating the effects of different architectural choices and using context-parallelism to enable training on sequences up to one minute (1M tokens) in length. We present a model, HarmonicRNN, which attains state of the art log-likelihoods and perceptual metrics on small-scale datasets.
title Linear RNNs for autoregressive generation of long music samples
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2510.02401