CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Munakata, Hokuto, Imamura, Takehiro, Nishimura, Taichi, Komatsu, Tatsuya
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908795232321536
author Munakata, Hokuto
Imamura, Takehiro
Nishimura, Taichi
Komatsu, Tatsuya
author_facet Munakata, Hokuto
Imamura, Takehiro
Nishimura, Taichi
Komatsu, Tatsuya
contents We introduce CASTELLA, a human-annotated audio benchmark for the task of audio moment retrieval (AMR). Although AMR has various useful potential applications, there is still no established benchmark with real-world data. The initial study of AMR trained the models solely on synthetic datasets. Moreover, the evaluation is based on an annotated dataset of fewer than 100 samples. This resulted in less reliable reported performance. To ensure performance for applications in real-world environments, we present CASTELLA, a large-scale manually annotated AMR dataset. CASTELLA consists of 1009, 213, and 640 audio recordings for training, validation, and test splits, respectively, which is 24 times larger than the previous dataset. We also establish a baseline model for AMR using CASTELLA. Our experiments demonstrate that a model fine-tuned on CASTELLA after pre-training on the synthetic data outperformed a model trained solely on the synthetic data by 10.4 points in Recall1@0.7. CASTELLA is publicly available in https://h-munakata.github.io/CASTELLA-demo/.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15131
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries
Munakata, Hokuto
Imamura, Takehiro
Nishimura, Taichi
Komatsu, Tatsuya
Audio and Speech Processing
Computation and Language
Sound
We introduce CASTELLA, a human-annotated audio benchmark for the task of audio moment retrieval (AMR). Although AMR has various useful potential applications, there is still no established benchmark with real-world data. The initial study of AMR trained the models solely on synthetic datasets. Moreover, the evaluation is based on an annotated dataset of fewer than 100 samples. This resulted in less reliable reported performance. To ensure performance for applications in real-world environments, we present CASTELLA, a large-scale manually annotated AMR dataset. CASTELLA consists of 1009, 213, and 640 audio recordings for training, validation, and test splits, respectively, which is 24 times larger than the previous dataset. We also establish a baseline model for AMR using CASTELLA. Our experiments demonstrate that a model fine-tuned on CASTELLA after pre-training on the synthetic data outperformed a model trained solely on the synthetic data by 10.4 points in Recall1@0.7. CASTELLA is publicly available in https://h-munakata.github.io/CASTELLA-demo/.
title CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2511.15131