Saved in:
Bibliographic Details
Main Authors: Kim, Hyun Jun, Choi, Hyeong Yong, Lim, Changwon
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.16649
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916958991024128
author Kim, Hyun Jun
Choi, Hyeong Yong
Lim, Changwon
author_facet Kim, Hyun Jun
Choi, Hyeong Yong
Lim, Changwon
contents This report presents the AISTAT team's submission to the language-based audio retrieval task in DCASE 2025 Task 6. Our proposed system employs dual encoder architecture, where audio and text modalities are encoded separately, and their representations are aligned using contrastive learning. Drawing inspiration from methodologies of the previous year's challenge, we implemented a distillation approach and leveraged large language models (LLMs) for effective data augmentation techniques, including back-translation and LLM mix. Additionally, we incorporated clustering to introduce an auxiliary classification task for further finetuning. Our best single system achieved a mAP@16 of 46.62, while an ensemble of four systems reached a mAP@16 of 48.83 on the Clotho development test split.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16649
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AISTAT lab system for DCASE2025 Task6: Language-based audio retrieval
Kim, Hyun Jun
Choi, Hyeong Yong
Lim, Changwon
Sound
Artificial Intelligence
Audio and Speech Processing
This report presents the AISTAT team's submission to the language-based audio retrieval task in DCASE 2025 Task 6. Our proposed system employs dual encoder architecture, where audio and text modalities are encoded separately, and their representations are aligned using contrastive learning. Drawing inspiration from methodologies of the previous year's challenge, we implemented a distillation approach and leveraged large language models (LLMs) for effective data augmentation techniques, including back-translation and LLM mix. Additionally, we incorporated clustering to introduce an auxiliary classification task for further finetuning. Our best single system achieved a mAP@16 of 46.62, while an ensemble of four systems reached a mAP@16 of 48.83 on the Clotho development test split.
title AISTAT lab system for DCASE2025 Task6: Language-based audio retrieval
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2509.16649