AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yuanyuan, Chen, Hangting, Yang, Dongchao, Wu, Zhiyong, Wu, Xixin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913767476953088
author Wang, Yuanyuan
Chen, Hangting
Yang, Dongchao
Wu, Zhiyong
Wu, Xixin
author_facet Wang, Yuanyuan
Chen, Hangting
Yang, Dongchao
Wu, Zhiyong
Wu, Xixin
contents Current Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content and style. Some studies try to improve the granularity by incorporating additional frame-level conditions or control networks. However, this usually leads to complex system design and difficulties due to the requirement for reference frame-level conditions. To address these challenges, we propose AudioComposer, a novel TTA generation framework that relies solely on natural language descriptions (NLDs) to provide both content specification and style control information. To further enhance audio generative modeling, we employ flow-based diffusion transformers with the cross-attention mechanism to incorporate text descriptions effectively into audio generation processes, which can not only simultaneously consider the content and style information in the text inputs, but also accelerate generation compared to other architectures. Furthermore, we propose a novel and comprehensive automatic data simulation pipeline to construct data with fine-grained text descriptions, which significantly alleviates the problem of data scarcity in the area. Experiments demonstrate the effectiveness of our framework using solely NLDs as inputs for content specification and style control. The generation quality and controllability surpass state-of-the-art TTA models, even with a smaller model size.
format Preprint
id arxiv_https___arxiv_org_abs_2409_12560
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
Wang, Yuanyuan
Chen, Hangting
Yang, Dongchao
Wu, Zhiyong
Wu, Xixin
Audio and Speech Processing
Sound
Current Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content and style. Some studies try to improve the granularity by incorporating additional frame-level conditions or control networks. However, this usually leads to complex system design and difficulties due to the requirement for reference frame-level conditions. To address these challenges, we propose AudioComposer, a novel TTA generation framework that relies solely on natural language descriptions (NLDs) to provide both content specification and style control information. To further enhance audio generative modeling, we employ flow-based diffusion transformers with the cross-attention mechanism to incorporate text descriptions effectively into audio generation processes, which can not only simultaneously consider the content and style information in the text inputs, but also accelerate generation compared to other architectures. Furthermore, we propose a novel and comprehensive automatic data simulation pipeline to construct data with fine-grained text descriptions, which significantly alleviates the problem of data scarcity in the area. Experiments demonstrate the effectiveness of our framework using solely NLDs as inputs for content specification and style control. The generation quality and controllability surpass state-of-the-art TTA models, even with a smaller model size.
title AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2409.12560