Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yuan, Yi, Liu, Haohe, Liang, Jinhua, Liu, Xubo, Plumbley, Mark D., Wang, Wenwu
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913449430220800
author Yuan, Yi
Liu, Haohe
Liang, Jinhua
Liu, Xubo
Plumbley, Mark D.
Wang, Wenwu
author_facet Yuan, Yi
Liu, Haohe
Liang, Jinhua
Liu, Xubo
Plumbley, Mark D.
Wang, Wenwu
contents Deep neural networks have recently achieved breakthroughs in sound generation. Despite the outstanding sample quality, current sound generation models face issues on small-scale datasets (e.g., overfitting), significantly limiting performance. In this paper, we make the first attempt to investigate the benefits of pre-training on sound generation with AudioLDM, the cutting-edge model for audio generation, as the backbone. Our study demonstrates the advantages of the pre-trained AudioLDM, especially in data-scarcity scenarios. In addition, the baselines and evaluation protocol for sound generation systems are not consistent enough to compare different studies directly. Aiming to facilitate further study on sound generation tasks, we benchmark the sound generation task on various frequently-used datasets. We hope our results on transfer learning and benchmarks can provide references for further research on conditional sound generation.
format Preprint
id arxiv_https___arxiv_org_abs_2303_03857
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study
Yuan, Yi
Liu, Haohe
Liang, Jinhua
Liu, Xubo
Plumbley, Mark D.
Wang, Wenwu
Sound
Artificial Intelligence
Multimedia
Audio and Speech Processing
Deep neural networks have recently achieved breakthroughs in sound generation. Despite the outstanding sample quality, current sound generation models face issues on small-scale datasets (e.g., overfitting), significantly limiting performance. In this paper, we make the first attempt to investigate the benefits of pre-training on sound generation with AudioLDM, the cutting-edge model for audio generation, as the backbone. Our study demonstrates the advantages of the pre-trained AudioLDM, especially in data-scarcity scenarios. In addition, the baselines and evaluation protocol for sound generation systems are not consistent enough to compare different studies directly. Aiming to facilitate further study on sound generation tasks, we benchmark the sound generation task on various frequently-used datasets. We hope our results on transfer learning and benchmarks can provide references for further research on conditional sound generation.
title Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study
topic Sound
Artificial Intelligence
Multimedia
Audio and Speech Processing
url https://arxiv.org/abs/2303.03857