STAB: Speech Tokenizer Assessment Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vashishth, Shikhar, Singh, Harman, Bharadwaj, Shikhar, Ganapathy, Sriram, Asawaroengchai, Chulayuth, Audhkhasi, Kartik, Rosenberg, Andrew, Bapna, Ankur, Ramabhadran, Bhuvana
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912014225375232
author Vashishth, Shikhar
Singh, Harman
Bharadwaj, Shikhar
Ganapathy, Sriram
Asawaroengchai, Chulayuth
Audhkhasi, Kartik
Rosenberg, Andrew
Bapna, Ankur
Ramabhadran, Bhuvana
author_facet Vashishth, Shikhar
Singh, Harman
Bharadwaj, Shikhar
Ganapathy, Sriram
Asawaroengchai, Chulayuth
Audhkhasi, Kartik
Rosenberg, Andrew
Bapna, Ankur
Ramabhadran, Bhuvana
contents Representing speech as discrete tokens provides a framework for transforming speech into a format that closely resembles text, thus enabling the use of speech as an input to the widely successful large language models (LLMs). Currently, while several speech tokenizers have been proposed, there is ambiguity regarding the properties that are desired from a tokenizer for specific downstream tasks and its overall generalizability. Evaluating the performance of tokenizers across different downstream tasks is a computationally intensive effort that poses challenges for scalability. To circumvent this requirement, we present STAB (Speech Tokenizer Assessment Benchmark), a systematic evaluation framework designed to assess speech tokenizers comprehensively and shed light on their inherent characteristics. This framework provides a deeper understanding of the underlying mechanisms of speech tokenization, thereby offering a valuable resource for expediting the advancement of future tokenizer models and enabling comparative analysis using a standardized benchmark. We evaluate the STAB metrics and correlate this with downstream task performance across a range of speech tasks and tokenizer choices.
format Preprint
id arxiv_https___arxiv_org_abs_2409_02384
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle STAB: Speech Tokenizer Assessment Benchmark
Vashishth, Shikhar
Singh, Harman
Bharadwaj, Shikhar
Ganapathy, Sriram
Asawaroengchai, Chulayuth
Audhkhasi, Kartik
Rosenberg, Andrew
Bapna, Ankur
Ramabhadran, Bhuvana
Computation and Language
Sound
Audio and Speech Processing
Representing speech as discrete tokens provides a framework for transforming speech into a format that closely resembles text, thus enabling the use of speech as an input to the widely successful large language models (LLMs). Currently, while several speech tokenizers have been proposed, there is ambiguity regarding the properties that are desired from a tokenizer for specific downstream tasks and its overall generalizability. Evaluating the performance of tokenizers across different downstream tasks is a computationally intensive effort that poses challenges for scalability. To circumvent this requirement, we present STAB (Speech Tokenizer Assessment Benchmark), a systematic evaluation framework designed to assess speech tokenizers comprehensively and shed light on their inherent characteristics. This framework provides a deeper understanding of the underlying mechanisms of speech tokenization, thereby offering a valuable resource for expediting the advancement of future tokenizer models and enabling comparative analysis using a standardized benchmark. We evaluate the STAB metrics and correlate this with downstream task performance across a range of speech tasks and tokenizer choices.
title STAB: Speech Tokenizer Assessment Benchmark
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2409.02384