Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yuan, Yi, Jia, Dongya, Zhuang, Xiaobin, Chen, Yuanzhe, Liu, Zhengxi, Chen, Zhuo, Wang, Yuping, Wang, Yuxuan, Liu, Xubo, Kang, Xiyuan, Plumbley, Mark D., Wang, Wenwu
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913630819188736
author Yuan, Yi
Jia, Dongya
Zhuang, Xiaobin
Chen, Yuanzhe
Liu, Zhengxi
Chen, Zhuo
Wang, Yuping
Wang, Yuxuan
Liu, Xubo
Kang, Xiyuan
Plumbley, Mark D.
Wang, Wenwu
author_facet Yuan, Yi
Jia, Dongya
Zhuang, Xiaobin
Chen, Yuanzhe
Liu, Zhengxi
Chen, Zhuo
Wang, Yuping
Wang, Yuxuan
Liu, Xubo
Kang, Xiyuan
Plumbley, Mark D.
Wang, Wenwu
contents Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from the simplicity and scarcity of the training data. This work aims to create a large-scale audio dataset with rich captions for improving audio generation models. We first develop an automated pipeline to generate detailed captions by transforming predicted visual captions, audio captions, and tagging labels into comprehensive descriptions using a Large Language Model (LLM). The resulting dataset, Sound-VECaps, comprises 1.66M high-quality audio-caption pairs with enriched details including audio event orders, occurred places and environment information. We then demonstrate that training the text-to-audio generation models with Sound-VECaps significantly improves the performance on complex prompts. Furthermore, we conduct ablation studies of the models on several downstream audio-language tasks, showing the potential of Sound-VECaps in advancing audio-text representation learning. Our dataset and models are available online from here https://yyua8222.github.io/Sound-VECaps-demo/.
format Preprint
id arxiv_https___arxiv_org_abs_2407_04416
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
Yuan, Yi
Jia, Dongya
Zhuang, Xiaobin
Chen, Yuanzhe
Liu, Zhengxi
Chen, Zhuo
Wang, Yuping
Wang, Yuxuan
Liu, Xubo
Kang, Xiyuan
Plumbley, Mark D.
Wang, Wenwu
Sound
Multimedia
Audio and Speech Processing
Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from the simplicity and scarcity of the training data. This work aims to create a large-scale audio dataset with rich captions for improving audio generation models. We first develop an automated pipeline to generate detailed captions by transforming predicted visual captions, audio captions, and tagging labels into comprehensive descriptions using a Large Language Model (LLM). The resulting dataset, Sound-VECaps, comprises 1.66M high-quality audio-caption pairs with enriched details including audio event orders, occurred places and environment information. We then demonstrate that training the text-to-audio generation models with Sound-VECaps significantly improves the performance on complex prompts. Furthermore, we conduct ablation studies of the models on several downstream audio-language tasks, showing the potential of Sound-VECaps in advancing audio-text representation learning. Our dataset and models are available online from here https://yyua8222.github.io/Sound-VECaps-demo/.
title Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
topic Sound
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2407.04416