Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kando, Shunsuke, Miyao, Yusuke, Takamichi, Shinnosuke
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918040996675584
author Kando, Shunsuke
Miyao, Yusuke
Takamichi, Shinnosuke
author_facet Kando, Shunsuke
Miyao, Yusuke
Takamichi, Shinnosuke
contents The purpose of speech tokenization is to transform a speech signal into a sequence of discrete representations, serving as the foundation for speech language models (SLMs). While speech tokenization has many options, their effect on the performance of SLMs remains unclear. This paper investigates two key aspects of speech tokenization: the segmentation width and the cluster size of discrete units. First, we segment speech signals into fixed/variable widths and pooled representations. We then train K-means models in multiple cluster sizes. Through the evaluation on zero-shot spoken language understanding benchmarks, we find the positive effect of moderately coarse segmentation and bigger cluster size. Notably, among the best-performing models, the most efficient one achieves a 50% reduction in training data and a 70% decrease in training runtime. Our analysis highlights the importance of combining multiple tokens to enhance fine-grained spoken language understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17446
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
Kando, Shunsuke
Miyao, Yusuke
Takamichi, Shinnosuke
Computation and Language
Sound
Audio and Speech Processing
The purpose of speech tokenization is to transform a speech signal into a sequence of discrete representations, serving as the foundation for speech language models (SLMs). While speech tokenization has many options, their effect on the performance of SLMs remains unclear. This paper investigates two key aspects of speech tokenization: the segmentation width and the cluster size of discrete units. First, we segment speech signals into fixed/variable widths and pooled representations. We then train K-means models in multiple cluster sizes. Through the evaluation on zero-shot spoken language understanding benchmarks, we find the positive effect of moderately coarse segmentation and bigger cluster size. Notably, among the best-performing models, the most efficient one achieves a 50% reduction in training data and a 70% decrease in training runtime. Our analysis highlights the importance of combining multiple tokens to enhance fine-grained spoken language understanding.
title Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
topic Computation and Language
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2505.17446