Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Luoyi, Xu, Xuenan, Wu, Mengyue, Xie, Weidi
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929490943279104
author Sun, Luoyi
Xu, Xuenan
Wu, Mengyue
Xie, Weidi
author_facet Sun, Luoyi
Xu, Xuenan
Wu, Mengyue
Xie, Weidi
contents Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the following aspects: insufficient volume, simplistic content, and arduous collection procedures. To establish an audio dataset with high-quality captions, we propose an innovative, automatic approach leveraging multimodal inputs, such as video frames, audio streams. Specifically, we construct a large-scale, high-quality, audio-language dataset, named as Auto-ACD, comprising over 1.5M audio-text pairs. We exploit a series of pre-trained models or APIs, to determine audio-visual synchronisation, generate image captions, object detection, or audio tags for specific videos. Subsequently, we employ LLM to paraphrase a congruent caption for each audio, guided by the extracted multi-modality clues. To demonstrate the effectiveness of the proposed dataset, we train widely used models on our dataset and show performance improvement on various downstream tasks, for example, audio-language retrieval, audio captioning, zero-shot classification. In addition, we establish a novel benchmark with environmental information and provide a benchmark for audio-text tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2309_11500
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning
Sun, Luoyi
Xu, Xuenan
Wu, Mengyue
Xie, Weidi
Sound
Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the following aspects: insufficient volume, simplistic content, and arduous collection procedures. To establish an audio dataset with high-quality captions, we propose an innovative, automatic approach leveraging multimodal inputs, such as video frames, audio streams. Specifically, we construct a large-scale, high-quality, audio-language dataset, named as Auto-ACD, comprising over 1.5M audio-text pairs. We exploit a series of pre-trained models or APIs, to determine audio-visual synchronisation, generate image captions, object detection, or audio tags for specific videos. Subsequently, we employ LLM to paraphrase a congruent caption for each audio, guided by the extracted multi-modality clues. To demonstrate the effectiveness of the proposed dataset, we train widely used models on our dataset and show performance improvement on various downstream tasks, for example, audio-language retrieval, audio captioning, zero-shot classification. In addition, we establish a novel benchmark with environmental information and provide a benchmark for audio-text tasks.
title Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning
topic Sound
Computer Vision and Pattern Recognition
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2309.11500