Discrete Audio Tokens: More Than a Survey!

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mousavi, Pooneh, Maimon, Gallil, Moumen, Adel, Petermann, Darius, Shi, Jiatong, Wu, Haibin, Yang, Haici, Kuznetsova, Anastasia, Ploujnikov, Artem, Marxer, Ricard, Ramabhadran, Bhuvana, Elizalde, Benjamin, Lugosch, Loren, Li, Jinyu, Subakan, Cem, Woodland, Phil, Kim, Minje, Lee, Hung-yi, Watanabe, Shinji, Adi, Yossi, Ravanelli, Mirco
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912610923839488
author Mousavi, Pooneh
Maimon, Gallil
Moumen, Adel
Petermann, Darius
Shi, Jiatong
Wu, Haibin
Yang, Haici
Kuznetsova, Anastasia
Ploujnikov, Artem
Marxer, Ricard
Ramabhadran, Bhuvana
Elizalde, Benjamin
Lugosch, Loren
Li, Jinyu
Subakan, Cem
Woodland, Phil
Kim, Minje
Lee, Hung-yi
Watanabe, Shinji
Adi, Yossi
Ravanelli, Mirco
author_facet Mousavi, Pooneh
Maimon, Gallil
Moumen, Adel
Petermann, Darius
Shi, Jiatong
Wu, Haibin
Yang, Haici
Kuznetsova, Anastasia
Ploujnikov, Artem
Marxer, Ricard
Ramabhadran, Bhuvana
Elizalde, Benjamin
Lugosch, Loren
Li, Jinyu
Subakan, Cem
Woodland, Phil
Kim, Minje
Lee, Hung-yi
Watanabe, Shinji
Adi, Yossi
Ravanelli, Mirco
contents Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and inference, as well as competitive performance across diverse downstream tasks. They provide a practical alternative to continuous features, enabling the integration of speech and audio into modern large language models (LLMs). As interest in token-based audio processing grows, various tokenization methods have emerged, and several surveys have reviewed the latest progress in the field. However, existing studies often focus on specific domains or tasks and lack a unified comparison across various benchmarks. This paper presents a systematic review and benchmark of discrete audio tokenizers, covering three domains: speech, music, and general audio. We propose a taxonomy of tokenization approaches based on encoder-decoder, quantization techniques, training paradigm, streamability, and application domains. We evaluate tokenizers on multiple benchmarks for reconstruction, downstream performance, and acoustic language modeling, and analyze trade-offs through controlled ablation studies. Our findings highlight key limitations, practical considerations, and open challenges, providing insight and guidance for future research in this rapidly evolving area. For more information, including our main results and tokenizer database, please refer to our website: https://poonehmousavi.github.io/dates-website/.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10274
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Discrete Audio Tokens: More Than a Survey!
Mousavi, Pooneh
Maimon, Gallil
Moumen, Adel
Petermann, Darius
Shi, Jiatong
Wu, Haibin
Yang, Haici
Kuznetsova, Anastasia
Ploujnikov, Artem
Marxer, Ricard
Ramabhadran, Bhuvana
Elizalde, Benjamin
Lugosch, Loren
Li, Jinyu
Subakan, Cem
Woodland, Phil
Kim, Minje
Lee, Hung-yi
Watanabe, Shinji
Adi, Yossi
Ravanelli, Mirco
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and inference, as well as competitive performance across diverse downstream tasks. They provide a practical alternative to continuous features, enabling the integration of speech and audio into modern large language models (LLMs). As interest in token-based audio processing grows, various tokenization methods have emerged, and several surveys have reviewed the latest progress in the field. However, existing studies often focus on specific domains or tasks and lack a unified comparison across various benchmarks. This paper presents a systematic review and benchmark of discrete audio tokenizers, covering three domains: speech, music, and general audio. We propose a taxonomy of tokenization approaches based on encoder-decoder, quantization techniques, training paradigm, streamability, and application domains. We evaluate tokenizers on multiple benchmarks for reconstruction, downstream performance, and acoustic language modeling, and analyze trade-offs through controlled ablation studies. Our findings highlight key limitations, practical considerations, and open challenges, providing insight and guidance for future research in this rapidly evolving area. For more information, including our main results and tokenizer database, please refer to our website: https://poonehmousavi.github.io/dates-website/.
title Discrete Audio Tokens: More Than a Survey!
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2506.10274