EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912438991978496 |
|---|---|
| author | Hai, Jiarui Xu, Yong Zhang, Hao Li, Chenxing Wang, Helin Elhilali, Mounya Yu, Dong |
| author_facet | Hai, Jiarui Xu, Yong Zhang, Hao Li, Chenxing Wang, Helin Elhilali, Mounya Yu, Dong |
| contents | We introduce EzAudio, a text-to-audio (T2A) generation framework designed to produce high-quality, natural-sounding sound effects. Core designs include: (1) We propose EzAudio-DiT, an optimized Diffusion Transformer (DiT) designed for audio latent representations, improving convergence speed, as well as parameter and memory efficiency. (2) We apply a classifier-free guidance (CFG) rescaling technique to mitigate fidelity loss at higher CFG scores and enhancing prompt adherence without compromising audio quality. (3) We propose a synthetic caption generation strategy leveraging recent advances in audio understanding and LLMs to enhance T2A pretraining. We show that EzAudio, with its computationally efficient architecture and fast convergence, is a competitive open-source model that excels in both objective and subjective evaluations by delivering highly realistic listening experiences. Code, data, and pre-trained models are released at: https://haidog-yaqub.github.io/EzAudio-Page/. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_10819 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer Hai, Jiarui Xu, Yong Zhang, Hao Li, Chenxing Wang, Helin Elhilali, Mounya Yu, Dong Audio and Speech Processing Sound We introduce EzAudio, a text-to-audio (T2A) generation framework designed to produce high-quality, natural-sounding sound effects. Core designs include: (1) We propose EzAudio-DiT, an optimized Diffusion Transformer (DiT) designed for audio latent representations, improving convergence speed, as well as parameter and memory efficiency. (2) We apply a classifier-free guidance (CFG) rescaling technique to mitigate fidelity loss at higher CFG scores and enhancing prompt adherence without compromising audio quality. (3) We propose a synthetic caption generation strategy leveraging recent advances in audio understanding and LLMs to enhance T2A pretraining. We show that EzAudio, with its computationally efficient architecture and fast convergence, is a competitive open-source model that excels in both objective and subjective evaluations by delivering highly realistic listening experiences. Code, data, and pre-trained models are released at: https://haidog-yaqub.github.io/EzAudio-Page/. |
| title | EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer |
| topic | Audio and Speech Processing Sound |
| url | https://arxiv.org/abs/2409.10819 |