DITTO: Diffusion Inference-Time T-Optimization for Music Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Novack, Zachary, McAuley, Julian, Berg-Kirkpatrick, Taylor, Bryan, Nicholas J.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914820128768000
author Novack, Zachary
McAuley, Julian
Berg-Kirkpatrick, Taylor
Bryan, Nicholas J.
author_facet Novack, Zachary
McAuley, Julian
Berg-Kirkpatrick, Taylor
Bryan, Nicholas J.
contents We propose Diffusion Inference-Time T-Optimization (DITTO), a general-purpose frame-work for controlling pre-trained text-to-music diffusion models at inference-time via optimizing initial noise latents. Our method can be used to optimize through any differentiable feature matching loss to achieve a target (stylized) output and leverages gradient checkpointing for memory efficiency. We demonstrate a surprisingly wide-range of applications for music generation including inpainting, outpainting, and looping as well as intensity, melody, and musical structure control - all without ever fine-tuning the underlying model. When we compare our approach against related training, guidance, and optimization-based methods, we find DITTO achieves state-of-the-art performance on nearly all tasks, including outperforming comparable approaches on controllability, audio quality, and computational efficiency, thus opening the door for high-quality, flexible, training-free control of diffusion models. Sound examples can be found at https://DITTO-Music.github.io/web/.
format Preprint
id arxiv_https___arxiv_org_abs_2401_12179
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DITTO: Diffusion Inference-Time T-Optimization for Music Generation
Novack, Zachary
McAuley, Julian
Berg-Kirkpatrick, Taylor
Bryan, Nicholas J.
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
We propose Diffusion Inference-Time T-Optimization (DITTO), a general-purpose frame-work for controlling pre-trained text-to-music diffusion models at inference-time via optimizing initial noise latents. Our method can be used to optimize through any differentiable feature matching loss to achieve a target (stylized) output and leverages gradient checkpointing for memory efficiency. We demonstrate a surprisingly wide-range of applications for music generation including inpainting, outpainting, and looping as well as intensity, melody, and musical structure control - all without ever fine-tuning the underlying model. When we compare our approach against related training, guidance, and optimization-based methods, we find DITTO achieves state-of-the-art performance on nearly all tasks, including outperforming comparable approaches on controllability, audio quality, and computational efficiency, thus opening the door for high-quality, flexible, training-free control of diffusion models. Sound examples can be found at https://DITTO-Music.github.io/web/.
title DITTO: Diffusion Inference-Time T-Optimization for Music Generation
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2401.12179