Continuous Autoregressive Models with Noise Augmentation Avoid Error Accumulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pasini, Marco, Nistal, Javier, Lattner, Stefan, Fazekas, George
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915037392666624
author Pasini, Marco
Nistal, Javier
Lattner, Stefan
Fazekas, George
author_facet Pasini, Marco
Nistal, Javier
Lattner, Stefan
Fazekas, George
contents Autoregressive models are typically applied to sequences of discrete tokens, but recent research indicates that generating sequences of continuous embeddings in an autoregressive manner is also feasible. However, such Continuous Autoregressive Models (CAMs) can suffer from a decline in generation quality over extended sequences due to error accumulation during inference. We introduce a novel method to address this issue by injecting random noise into the input embeddings during training. This procedure makes the model robust against varying error levels at inference. We further reduce error accumulation through an inference procedure that introduces low-level noise. Experiments on musical audio generation show that CAM substantially outperforms existing autoregressive and non-autoregressive approaches while preserving audio quality over extended sequences. This work paves the way for generating continuous embeddings in a purely autoregressive setting, opening new possibilities for real-time and interactive generative applications.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18447
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Continuous Autoregressive Models with Noise Augmentation Avoid Error Accumulation
Pasini, Marco
Nistal, Javier
Lattner, Stefan
Fazekas, George
Machine Learning
Artificial Intelligence
Sound
Audio and Speech Processing
Autoregressive models are typically applied to sequences of discrete tokens, but recent research indicates that generating sequences of continuous embeddings in an autoregressive manner is also feasible. However, such Continuous Autoregressive Models (CAMs) can suffer from a decline in generation quality over extended sequences due to error accumulation during inference. We introduce a novel method to address this issue by injecting random noise into the input embeddings during training. This procedure makes the model robust against varying error levels at inference. We further reduce error accumulation through an inference procedure that introduces low-level noise. Experiments on musical audio generation show that CAM substantially outperforms existing autoregressive and non-autoregressive approaches while preserving audio quality over extended sequences. This work paves the way for generating continuous embeddings in a purely autoregressive setting, opening new possibilities for real-time and interactive generative applications.
title Continuous Autoregressive Models with Noise Augmentation Avoid Error Accumulation
topic Machine Learning
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2411.18447