Presto! Distilling Steps and Layers for Accelerating Music Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866912331988992000 |
|---|---|
| author | Novack, Zachary Zhu, Ge Casebeer, Jonah McAuley, Julian Berg-Kirkpatrick, Taylor Bryan, Nicholas J. |
| author_facet | Novack, Zachary Zhu, Ge Casebeer, Jonah McAuley, Julian Berg-Kirkpatrick, Taylor Bryan, Nicholas J. |
| contents | Despite advances in diffusion-based text-to-music (TTM) methods, efficient, high-quality generation remains a challenge. We introduce Presto!, an approach to inference acceleration for score-based diffusion transformers via reducing both sampling steps and cost per step. To reduce steps, we develop a new score-based distribution matching distillation (DMD) method for the EDM-family of diffusion models, the first GAN-based distillation method for TTM. To reduce the cost per step, we develop a simple, but powerful improvement to a recent layer distillation method that improves learning via better preserving hidden state variance. Finally, we combine our step and layer distillation methods together for a dual-faceted approach. We evaluate our step and layer distillation methods independently and show each yield best-in-class performance. Our combined distillation method can generate high-quality outputs with improved diversity, accelerating our base model by 10-18x (230/435ms latency for 32 second mono/stereo 44.1kHz, 15x faster than comparable SOTA) -- the fastest high-quality TTM to our knowledge. Sound examples can be found at https://presto-music.github.io/web/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_05167 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Presto! Distilling Steps and Layers for Accelerating Music Generation Novack, Zachary Zhu, Ge Casebeer, Jonah McAuley, Julian Berg-Kirkpatrick, Taylor Bryan, Nicholas J. Sound Artificial Intelligence Machine Learning Audio and Speech Processing Despite advances in diffusion-based text-to-music (TTM) methods, efficient, high-quality generation remains a challenge. We introduce Presto!, an approach to inference acceleration for score-based diffusion transformers via reducing both sampling steps and cost per step. To reduce steps, we develop a new score-based distribution matching distillation (DMD) method for the EDM-family of diffusion models, the first GAN-based distillation method for TTM. To reduce the cost per step, we develop a simple, but powerful improvement to a recent layer distillation method that improves learning via better preserving hidden state variance. Finally, we combine our step and layer distillation methods together for a dual-faceted approach. We evaluate our step and layer distillation methods independently and show each yield best-in-class performance. Our combined distillation method can generate high-quality outputs with improved diversity, accelerating our base model by 10-18x (230/435ms latency for 32 second mono/stereo 44.1kHz, 15x faster than comparable SOTA) -- the fastest high-quality TTM to our knowledge. Sound examples can be found at https://presto-music.github.io/web/. |
| title | Presto! Distilling Steps and Layers for Accelerating Music Generation |
| topic | Sound Artificial Intelligence Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2410.05167 |