EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hai, Jiarui, Xu, Yong, Zhang, Hao, Li, Chenxing, Wang, Helin, Elhilali, Mounya, Yu, Dong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912438991978496
author Hai, Jiarui
Xu, Yong
Zhang, Hao
Li, Chenxing
Wang, Helin
Elhilali, Mounya
Yu, Dong
author_facet Hai, Jiarui
Xu, Yong
Zhang, Hao
Li, Chenxing
Wang, Helin
Elhilali, Mounya
Yu, Dong
contents We introduce EzAudio, a text-to-audio (T2A) generation framework designed to produce high-quality, natural-sounding sound effects. Core designs include: (1) We propose EzAudio-DiT, an optimized Diffusion Transformer (DiT) designed for audio latent representations, improving convergence speed, as well as parameter and memory efficiency. (2) We apply a classifier-free guidance (CFG) rescaling technique to mitigate fidelity loss at higher CFG scores and enhancing prompt adherence without compromising audio quality. (3) We propose a synthetic caption generation strategy leveraging recent advances in audio understanding and LLMs to enhance T2A pretraining. We show that EzAudio, with its computationally efficient architecture and fast convergence, is a competitive open-source model that excels in both objective and subjective evaluations by delivering highly realistic listening experiences. Code, data, and pre-trained models are released at: https://haidog-yaqub.github.io/EzAudio-Page/.
format Preprint
id arxiv_https___arxiv_org_abs_2409_10819
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
Hai, Jiarui
Xu, Yong
Zhang, Hao
Li, Chenxing
Wang, Helin
Elhilali, Mounya
Yu, Dong
Audio and Speech Processing
Sound
We introduce EzAudio, a text-to-audio (T2A) generation framework designed to produce high-quality, natural-sounding sound effects. Core designs include: (1) We propose EzAudio-DiT, an optimized Diffusion Transformer (DiT) designed for audio latent representations, improving convergence speed, as well as parameter and memory efficiency. (2) We apply a classifier-free guidance (CFG) rescaling technique to mitigate fidelity loss at higher CFG scores and enhancing prompt adherence without compromising audio quality. (3) We propose a synthetic caption generation strategy leveraging recent advances in audio understanding and LLMs to enhance T2A pretraining. We show that EzAudio, with its computationally efficient architecture and fast convergence, is a competitive open-source model that excels in both objective and subjective evaluations by delivering highly realistic listening experiences. Code, data, and pre-trained models are released at: https://haidog-yaqub.github.io/EzAudio-Page/.
title EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2409.10819