AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhao, Junqi, Zhao, Jinzheng, Liu, Haohe, Chen, Yun, Han, Lu, Liu, Xubo, Plumbley, Mark, Wang, Wenwu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908382715183104
author Zhao, Junqi
Zhao, Jinzheng
Liu, Haohe
Chen, Yun
Han, Lu
Liu, Xubo
Plumbley, Mark
Wang, Wenwu
author_facet Zhao, Junqi
Zhao, Jinzheng
Liu, Haohe
Chen, Yun
Han, Lu
Liu, Xubo
Plumbley, Mark
Wang, Wenwu
contents Diffusion models have significantly improved the quality and diversity of audio generation but are hindered by slow inference speed. Rectified flow enhances inference speed by learning straight-line ordinary differential equation (ODE) paths. However, this approach requires training a flow-matching model from scratch and tends to perform suboptimally, or even poorly, at low step counts. To address the limitations of rectified flow while leveraging the advantages of advanced pre-trained diffusion models, this study integrates pre-trained models with the rectified diffusion method to improve the efficiency of text-to-audio (TTA) generation. Specifically, we propose AudioTurbo, which learns first-order ODE paths from deterministic noise sample pairs generated by a pre-trained TTA model. Experiments on the AudioCaps dataset demonstrate that our model, with only 10 sampling steps, outperforms prior models and reduces inference to 3 steps compared to a flow-matching-based acceleration model.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22106
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
Zhao, Junqi
Zhao, Jinzheng
Liu, Haohe
Chen, Yun
Han, Lu
Liu, Xubo
Plumbley, Mark
Wang, Wenwu
Sound
Artificial Intelligence
Audio and Speech Processing
Diffusion models have significantly improved the quality and diversity of audio generation but are hindered by slow inference speed. Rectified flow enhances inference speed by learning straight-line ordinary differential equation (ODE) paths. However, this approach requires training a flow-matching model from scratch and tends to perform suboptimally, or even poorly, at low step counts. To address the limitations of rectified flow while leveraging the advantages of advanced pre-trained diffusion models, this study integrates pre-trained models with the rectified diffusion method to improve the efficiency of text-to-audio (TTA) generation. Specifically, we propose AudioTurbo, which learns first-order ODE paths from deterministic noise sample pairs generated by a pre-trained TTA model. Experiments on the AudioCaps dataset demonstrate that our model, with only 10 sampling steps, outperforms prior models and reduces inference to 3 steps compared to a flow-matching-based acceleration model.
title AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2505.22106