Presto! Distilling Steps and Layers for Accelerating Music Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Novack, Zachary, Zhu, Ge, Casebeer, Jonah, McAuley, Julian, Berg-Kirkpatrick, Taylor, Bryan, Nicholas J.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912331988992000
author Novack, Zachary
Zhu, Ge
Casebeer, Jonah
McAuley, Julian
Berg-Kirkpatrick, Taylor
Bryan, Nicholas J.
author_facet Novack, Zachary
Zhu, Ge
Casebeer, Jonah
McAuley, Julian
Berg-Kirkpatrick, Taylor
Bryan, Nicholas J.
contents Despite advances in diffusion-based text-to-music (TTM) methods, efficient, high-quality generation remains a challenge. We introduce Presto!, an approach to inference acceleration for score-based diffusion transformers via reducing both sampling steps and cost per step. To reduce steps, we develop a new score-based distribution matching distillation (DMD) method for the EDM-family of diffusion models, the first GAN-based distillation method for TTM. To reduce the cost per step, we develop a simple, but powerful improvement to a recent layer distillation method that improves learning via better preserving hidden state variance. Finally, we combine our step and layer distillation methods together for a dual-faceted approach. We evaluate our step and layer distillation methods independently and show each yield best-in-class performance. Our combined distillation method can generate high-quality outputs with improved diversity, accelerating our base model by 10-18x (230/435ms latency for 32 second mono/stereo 44.1kHz, 15x faster than comparable SOTA) -- the fastest high-quality TTM to our knowledge. Sound examples can be found at https://presto-music.github.io/web/.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05167
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Presto! Distilling Steps and Layers for Accelerating Music Generation
Novack, Zachary
Zhu, Ge
Casebeer, Jonah
McAuley, Julian
Berg-Kirkpatrick, Taylor
Bryan, Nicholas J.
Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
Despite advances in diffusion-based text-to-music (TTM) methods, efficient, high-quality generation remains a challenge. We introduce Presto!, an approach to inference acceleration for score-based diffusion transformers via reducing both sampling steps and cost per step. To reduce steps, we develop a new score-based distribution matching distillation (DMD) method for the EDM-family of diffusion models, the first GAN-based distillation method for TTM. To reduce the cost per step, we develop a simple, but powerful improvement to a recent layer distillation method that improves learning via better preserving hidden state variance. Finally, we combine our step and layer distillation methods together for a dual-faceted approach. We evaluate our step and layer distillation methods independently and show each yield best-in-class performance. Our combined distillation method can generate high-quality outputs with improved diversity, accelerating our base model by 10-18x (230/435ms latency for 32 second mono/stereo 44.1kHz, 15x faster than comparable SOTA) -- the fastest high-quality TTM to our knowledge. Sound examples can be found at https://presto-music.github.io/web/.
title Presto! Distilling Steps and Layers for Accelerating Music Generation
topic Sound
Artificial Intelligence
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2410.05167