Long-form music generation with latent diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909271521755136 |
|---|---|
| author | Evans, Zach Parker, Julian D. Carr, CJ Zukowski, Zack Taylor, Josiah Pons, Jordi |
| author_facet | Evans, Zach Parker, Julian D. Carr, CJ Zukowski, Zack Taylor, Josiah Pons, Jordi |
| contents | Audio-based generative models for music have seen great strides recently, but so far have not managed to produce full-length music tracks with coherent musical structure from text prompts. We show that by training a generative model on long temporal contexts it is possible to produce long-form music of up to 4m45s. Our model consists of a diffusion-transformer operating on a highly downsampled continuous latent representation (latent rate of 21.5Hz). It obtains state-of-the-art generations according to metrics on audio quality and prompt alignment, and subjective tests reveal that it produces full-length music with coherent structure. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_10301 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Long-form music generation with latent diffusion Evans, Zach Parker, Julian D. Carr, CJ Zukowski, Zack Taylor, Josiah Pons, Jordi Sound Machine Learning Audio and Speech Processing Audio-based generative models for music have seen great strides recently, but so far have not managed to produce full-length music tracks with coherent musical structure from text prompts. We show that by training a generative model on long temporal contexts it is possible to produce long-form music of up to 4m45s. Our model consists of a diffusion-transformer operating on a highly downsampled continuous latent representation (latent rate of 21.5Hz). It obtains state-of-the-art generations according to metrics on audio quality and prompt alignment, and subjective tests reveal that it produces full-length music with coherent structure. |
| title | Long-form music generation with latent diffusion |
| topic | Sound Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2404.10301 |