ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914845795811328 |
|---|---|
| author | Bai, Yatong Dang, Trung Tran, Dung Koishida, Kazuhito Sojoudi, Somayeh |
| author_facet | Bai, Yatong Dang, Trung Tran, Dung Koishida, Kazuhito Sojoudi, Somayeh |
| contents | Diffusion models are instrumental in text-to-audio (TTA) generation. Unfortunately, they suffer from slow inference due to an excessive number of queries to the underlying denoising network per generation. To address this bottleneck, we introduce ConsistencyTTA, a framework requiring only a single non-autoregressive network query, thereby accelerating TTA by hundreds of times. We achieve so by proposing "CFG-aware latent consistency model," which adapts consistency generation into a latent space and incorporates classifier-free guidance (CFG) into model training. Moreover, unlike diffusion models, ConsistencyTTA can be finetuned closed-loop with audio-space text-aware metrics, such as CLAP score, to further enhance the generations. Our objective and subjective evaluation on the AudioCaps dataset shows that compared to diffusion-based counterparts, ConsistencyTTA reduces inference computation by 400x while retaining generation quality and diversity. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2309_10740 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation Bai, Yatong Dang, Trung Tran, Dung Koishida, Kazuhito Sojoudi, Somayeh Sound Machine Learning Multimedia Audio and Speech Processing Diffusion models are instrumental in text-to-audio (TTA) generation. Unfortunately, they suffer from slow inference due to an excessive number of queries to the underlying denoising network per generation. To address this bottleneck, we introduce ConsistencyTTA, a framework requiring only a single non-autoregressive network query, thereby accelerating TTA by hundreds of times. We achieve so by proposing "CFG-aware latent consistency model," which adapts consistency generation into a latent space and incorporates classifier-free guidance (CFG) into model training. Moreover, unlike diffusion models, ConsistencyTTA can be finetuned closed-loop with audio-space text-aware metrics, such as CLAP score, to further enhance the generations. Our objective and subjective evaluation on the AudioCaps dataset shows that compared to diffusion-based counterparts, ConsistencyTTA reduces inference computation by 400x while retaining generation quality and diversity. |
| title | ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation |
| topic | Sound Machine Learning Multimedia Audio and Speech Processing |
| url | https://arxiv.org/abs/2309.10740 |