Long-form music generation with latent diffusion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Evans, Zach, Parker, Julian D., Carr, CJ, Zukowski, Zack, Taylor, Josiah, Pons, Jordi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909271521755136
author Evans, Zach
Parker, Julian D.
Carr, CJ
Zukowski, Zack
Taylor, Josiah
Pons, Jordi
author_facet Evans, Zach
Parker, Julian D.
Carr, CJ
Zukowski, Zack
Taylor, Josiah
Pons, Jordi
contents Audio-based generative models for music have seen great strides recently, but so far have not managed to produce full-length music tracks with coherent musical structure from text prompts. We show that by training a generative model on long temporal contexts it is possible to produce long-form music of up to 4m45s. Our model consists of a diffusion-transformer operating on a highly downsampled continuous latent representation (latent rate of 21.5Hz). It obtains state-of-the-art generations according to metrics on audio quality and prompt alignment, and subjective tests reveal that it produces full-length music with coherent structure.
format Preprint
id arxiv_https___arxiv_org_abs_2404_10301
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Long-form music generation with latent diffusion
Evans, Zach
Parker, Julian D.
Carr, CJ
Zukowski, Zack
Taylor, Josiah
Pons, Jordi
Sound
Machine Learning
Audio and Speech Processing
Audio-based generative models for music have seen great strides recently, but so far have not managed to produce full-length music tracks with coherent musical structure from text prompts. We show that by training a generative model on long temporal contexts it is possible to produce long-form music of up to 4m45s. Our model consists of a diffusion-transformer operating on a highly downsampled continuous latent representation (latent rate of 21.5Hz). It obtains state-of-the-art generations according to metrics on audio quality and prompt alignment, and subjective tests reveal that it produces full-length music with coherent structure.
title Long-form music generation with latent diffusion
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2404.10301