EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Jaeyeon, Jeon, Minjeon, Jung, Jaeyoon, Woo, Sang Hoon, Lee, Jinjoo
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914933898215424
author Kim, Jaeyeon
Jeon, Minjeon
Jung, Jaeyoon
Woo, Sang Hoon
Lee, Jinjoo
author_facet Kim, Jaeyeon
Jeon, Minjeon
Jung, Jaeyoon
Woo, Sang Hoon
Lee, Jinjoo
contents In this work, we aim to analyze and optimize the EnCLAP framework, a state-of-the-art model in automated audio captioning. We investigate the impact of modifying the acoustic encoder components, explore pretraining with different dataset scales, and study the effectiveness of a reranking scheme. Through extensive experimentation and quantitative analysis of generated captions, we develop EnCLAP++, an enhanced version that significantly surpasses the original.
format Preprint
id arxiv_https___arxiv_org_abs_2409_01201
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance
Kim, Jaeyeon
Jeon, Minjeon
Jung, Jaeyoon
Woo, Sang Hoon
Lee, Jinjoo
Audio and Speech Processing
Artificial Intelligence
Sound
In this work, we aim to analyze and optimize the EnCLAP framework, a state-of-the-art model in automated audio captioning. We investigate the impact of modifying the acoustic encoder components, explore pretraining with different dataset scales, and study the effectiveness of a reranking scheme. Through extensive experimentation and quantitative analysis of generated captions, we develop EnCLAP++, an enhanced version that significantly surpasses the original.
title EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance
topic Audio and Speech Processing
Artificial Intelligence
Sound
url https://arxiv.org/abs/2409.01201