Discrete Audio Representations for Automated Audio Captioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Jingguang, Sun, Haoqin, Hu, Xinhui, Xu, Xinkang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912385371996160
author Tian, Jingguang
Sun, Haoqin
Hu, Xinhui
Xu, Xinkang
author_facet Tian, Jingguang
Sun, Haoqin
Hu, Xinhui
Xu, Xinkang
contents Discrete audio representations, termed audio tokens, are broadly categorized into semantic and acoustic tokens, typically generated through unsupervised tokenization of continuous audio representations. However, their applicability to automated audio captioning (AAC) remains underexplored. This paper systematically investigates the viability of audio token-driven models for AAC through comparative analyses of various tokenization methods. Our findings reveal that audio tokenization leads to performance degradation in AAC models compared to those that directly utilize continuous audio representations. To address this issue, we introduce a supervised audio tokenizer trained with an audio tagging objective. Unlike unsupervised tokenizers, which lack explicit semantic understanding, the proposed tokenizer effectively captures audio event information. Experiments conducted on the Clotho dataset demonstrate that the proposed audio tokens outperform conventional audio tokens in the AAC task.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14989
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Discrete Audio Representations for Automated Audio Captioning
Tian, Jingguang
Sun, Haoqin
Hu, Xinhui
Xu, Xinkang
Sound
Audio and Speech Processing
Discrete audio representations, termed audio tokens, are broadly categorized into semantic and acoustic tokens, typically generated through unsupervised tokenization of continuous audio representations. However, their applicability to automated audio captioning (AAC) remains underexplored. This paper systematically investigates the viability of audio token-driven models for AAC through comparative analyses of various tokenization methods. Our findings reveal that audio tokenization leads to performance degradation in AAC models compared to those that directly utilize continuous audio representations. To address this issue, we introduce a supervised audio tokenizer trained with an audio tagging objective. Unlike unsupervised tokenizers, which lack explicit semantic understanding, the proposed tokenizer effectively captures audio event information. Experiments conducted on the Clotho dataset demonstrate that the proposed audio tokens outperform conventional audio tokens in the AAC task.
title Discrete Audio Representations for Automated Audio Captioning
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.14989