ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bai, Yatong, Dang, Trung, Tran, Dung, Koishida, Kazuhito, Sojoudi, Somayeh
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914845795811328
author Bai, Yatong
Dang, Trung
Tran, Dung
Koishida, Kazuhito
Sojoudi, Somayeh
author_facet Bai, Yatong
Dang, Trung
Tran, Dung
Koishida, Kazuhito
Sojoudi, Somayeh
contents Diffusion models are instrumental in text-to-audio (TTA) generation. Unfortunately, they suffer from slow inference due to an excessive number of queries to the underlying denoising network per generation. To address this bottleneck, we introduce ConsistencyTTA, a framework requiring only a single non-autoregressive network query, thereby accelerating TTA by hundreds of times. We achieve so by proposing "CFG-aware latent consistency model," which adapts consistency generation into a latent space and incorporates classifier-free guidance (CFG) into model training. Moreover, unlike diffusion models, ConsistencyTTA can be finetuned closed-loop with audio-space text-aware metrics, such as CLAP score, to further enhance the generations. Our objective and subjective evaluation on the AudioCaps dataset shows that compared to diffusion-based counterparts, ConsistencyTTA reduces inference computation by 400x while retaining generation quality and diversity.
format Preprint
id arxiv_https___arxiv_org_abs_2309_10740
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
Bai, Yatong
Dang, Trung
Tran, Dung
Koishida, Kazuhito
Sojoudi, Somayeh
Sound
Machine Learning
Multimedia
Audio and Speech Processing
Diffusion models are instrumental in text-to-audio (TTA) generation. Unfortunately, they suffer from slow inference due to an excessive number of queries to the underlying denoising network per generation. To address this bottleneck, we introduce ConsistencyTTA, a framework requiring only a single non-autoregressive network query, thereby accelerating TTA by hundreds of times. We achieve so by proposing "CFG-aware latent consistency model," which adapts consistency generation into a latent space and incorporates classifier-free guidance (CFG) into model training. Moreover, unlike diffusion models, ConsistencyTTA can be finetuned closed-loop with audio-space text-aware metrics, such as CLAP score, to further enhance the generations. Our objective and subjective evaluation on the AudioCaps dataset shows that compared to diffusion-based counterparts, ConsistencyTTA reduces inference computation by 400x while retaining generation quality and diversity.
title ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
topic Sound
Machine Learning
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2309.10740