DreamAudio: Customized Text-to-Audio Generation with Diffusion Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Yi, Liu, Xubo, Liu, Haohe, Kang, Xiyuan, Chen, Zhuo, Wang, Yuxuan, Plumbley, Mark D., Wang, Wenwu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913063506018304
author Yuan, Yi
Liu, Xubo
Liu, Haohe
Kang, Xiyuan
Chen, Zhuo
Wang, Yuxuan
Plumbley, Mark D.
Wang, Wenwu
author_facet Yuan, Yi
Liu, Xubo
Liu, Haohe
Kang, Xiyuan
Chen, Zhuo
Wang, Yuxuan
Plumbley, Mark D.
Wang, Wenwu
contents With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite producing high-quality outputs, existing text-to-audio models mainly aim to generate semantically aligned sound and fall short of controlling fine-grained acoustic characteristics of specific sounds. As a result, users who need specific sound content may find it difficult to generate the desired audio clips. In this paper, we present DreamAudio for customized text-to-audio generation (CTTA). Specifically, we introduce a new framework that is designed to enable the model to identify auditory information from user-provided reference concepts for audio generation. Given a few reference audio samples containing personalized audio events, our system can generate new audio samples that include these specific events. In addition, two types of datasets are developed for training and testing the proposed systems. The experiments show that DreamAudio generates audio samples that are highly consistent with the customized audio features and aligned well with the input text prompts. Furthermore, DreamAudio offers comparable performance in general text-to-audio tasks. We also provide a human-involved dataset containing audio events from real-world CTTA cases as the benchmark for customized generation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06027
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
Yuan, Yi
Liu, Xubo
Liu, Haohe
Kang, Xiyuan
Chen, Zhuo
Wang, Yuxuan
Plumbley, Mark D.
Wang, Wenwu
Sound
Artificial Intelligence
Audio and Speech Processing
With the development of large-scale diffusion-based and language-modeling-based generative models, impressive progress has been achieved in text-to-audio generation. Despite producing high-quality outputs, existing text-to-audio models mainly aim to generate semantically aligned sound and fall short of controlling fine-grained acoustic characteristics of specific sounds. As a result, users who need specific sound content may find it difficult to generate the desired audio clips. In this paper, we present DreamAudio for customized text-to-audio generation (CTTA). Specifically, we introduce a new framework that is designed to enable the model to identify auditory information from user-provided reference concepts for audio generation. Given a few reference audio samples containing personalized audio events, our system can generate new audio samples that include these specific events. In addition, two types of datasets are developed for training and testing the proposed systems. The experiments show that DreamAudio generates audio samples that are highly consistent with the customized audio features and aligned well with the input text prompts. Furthermore, DreamAudio offers comparable performance in general text-to-audio tasks. We also provide a human-involved dataset containing audio events from real-world CTTA cases as the benchmark for customized generation tasks.
title DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2509.06027