Continuous Autoregressive Models with Noise Augmentation Avoid Error Accumulation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915037392666624 |
|---|---|
| author | Pasini, Marco Nistal, Javier Lattner, Stefan Fazekas, George |
| author_facet | Pasini, Marco Nistal, Javier Lattner, Stefan Fazekas, George |
| contents | Autoregressive models are typically applied to sequences of discrete tokens, but recent research indicates that generating sequences of continuous embeddings in an autoregressive manner is also feasible. However, such Continuous Autoregressive Models (CAMs) can suffer from a decline in generation quality over extended sequences due to error accumulation during inference. We introduce a novel method to address this issue by injecting random noise into the input embeddings during training. This procedure makes the model robust against varying error levels at inference. We further reduce error accumulation through an inference procedure that introduces low-level noise. Experiments on musical audio generation show that CAM substantially outperforms existing autoregressive and non-autoregressive approaches while preserving audio quality over extended sequences. This work paves the way for generating continuous embeddings in a purely autoregressive setting, opening new possibilities for real-time and interactive generative applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2411_18447 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Continuous Autoregressive Models with Noise Augmentation Avoid Error Accumulation Pasini, Marco Nistal, Javier Lattner, Stefan Fazekas, George Machine Learning Artificial Intelligence Sound Audio and Speech Processing Autoregressive models are typically applied to sequences of discrete tokens, but recent research indicates that generating sequences of continuous embeddings in an autoregressive manner is also feasible. However, such Continuous Autoregressive Models (CAMs) can suffer from a decline in generation quality over extended sequences due to error accumulation during inference. We introduce a novel method to address this issue by injecting random noise into the input embeddings during training. This procedure makes the model robust against varying error levels at inference. We further reduce error accumulation through an inference procedure that introduces low-level noise. Experiments on musical audio generation show that CAM substantially outperforms existing autoregressive and non-autoregressive approaches while preserving audio quality over extended sequences. This work paves the way for generating continuous embeddings in a purely autoregressive setting, opening new possibilities for real-time and interactive generative applications. |
| title | Continuous Autoregressive Models with Noise Augmentation Avoid Error Accumulation |
| topic | Machine Learning Artificial Intelligence Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2411.18447 |