AudioLCM: Text-to-Audio Generation with Latent Consistency Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Huadai, Huang, Rongjie, Liu, Yang, Cao, Hengyuan, Wang, Jialei, Cheng, Xize, Zheng, Siqi, Zhao, Zhou
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910519445684224
author Liu, Huadai
Huang, Rongjie
Liu, Yang
Cao, Hengyuan
Wang, Jialei
Cheng, Xize
Zheng, Siqi
Zhao, Zhou
author_facet Liu, Huadai
Huang, Rongjie
Liu, Yang
Cao, Hengyuan
Wang, Jialei
Cheng, Xize
Zheng, Siqi
Zhao, Zhou
contents Recent advancements in Latent Diffusion Models (LDMs) have propelled them to the forefront of various generative tasks. However, their iterative sampling process poses a significant computational burden, resulting in slow generation speeds and limiting their application in text-to-audio generation deployment. In this work, we introduce AudioLCM, a novel consistency-based model tailored for efficient and high-quality text-to-audio generation. AudioLCM integrates Consistency Models into the generation process, facilitating rapid inference through a mapping from any point at any time step to the trajectory's initial point. To overcome the convergence issue inherent in LDMs with reduced sample iterations, we propose the Guided Latent Consistency Distillation with a multi-step Ordinary Differential Equation (ODE) solver. This innovation shortens the time schedule from thousands to dozens of steps while maintaining sample quality, thereby achieving fast convergence and high-quality generation. Furthermore, to optimize the performance of transformer-based neural network architectures, we integrate the advanced techniques pioneered by LLaMA into the foundational framework of transformers. This architecture supports stable and efficient training, ensuring robust performance in text-to-audio synthesis. Experimental results on text-to-sound generation and text-to-music synthesis tasks demonstrate that AudioLCM needs only 2 iterations to synthesize high-fidelity audios, while it maintains sample quality competitive with state-of-the-art models using hundreds of steps. AudioLCM enables a sampling speed of 333x faster than real-time on a single NVIDIA 4090Ti GPU, making generative models practically applicable to text-to-audio generation deployment. Our extensive preliminary analysis shows that each design in AudioLCM is effective.
format Preprint
id arxiv_https___arxiv_org_abs_2406_00356
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AudioLCM: Text-to-Audio Generation with Latent Consistency Models
Liu, Huadai
Huang, Rongjie
Liu, Yang
Cao, Hengyuan
Wang, Jialei
Cheng, Xize
Zheng, Siqi
Zhao, Zhou
Audio and Speech Processing
Sound
Recent advancements in Latent Diffusion Models (LDMs) have propelled them to the forefront of various generative tasks. However, their iterative sampling process poses a significant computational burden, resulting in slow generation speeds and limiting their application in text-to-audio generation deployment. In this work, we introduce AudioLCM, a novel consistency-based model tailored for efficient and high-quality text-to-audio generation. AudioLCM integrates Consistency Models into the generation process, facilitating rapid inference through a mapping from any point at any time step to the trajectory's initial point. To overcome the convergence issue inherent in LDMs with reduced sample iterations, we propose the Guided Latent Consistency Distillation with a multi-step Ordinary Differential Equation (ODE) solver. This innovation shortens the time schedule from thousands to dozens of steps while maintaining sample quality, thereby achieving fast convergence and high-quality generation. Furthermore, to optimize the performance of transformer-based neural network architectures, we integrate the advanced techniques pioneered by LLaMA into the foundational framework of transformers. This architecture supports stable and efficient training, ensuring robust performance in text-to-audio synthesis. Experimental results on text-to-sound generation and text-to-music synthesis tasks demonstrate that AudioLCM needs only 2 iterations to synthesize high-fidelity audios, while it maintains sample quality competitive with state-of-the-art models using hundreds of steps. AudioLCM enables a sampling speed of 333x faster than real-time on a single NVIDIA 4090Ti GPU, making generative models practically applicable to text-to-audio generation deployment. Our extensive preliminary analysis shows that each design in AudioLCM is effective.
title AudioLCM: Text-to-Audio Generation with Latent Consistency Models
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2406.00356