Scaling Spoken Language Models with Syllabic Speech Tokenization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Nicholas, Cho, Cheol Jun, Black, Alan W, Anumanchipalli, Gopala K.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918321783308288
author Lee, Nicholas
Cho, Cheol Jun
Black, Alan W
Anumanchipalli, Gopala K.
author_facet Lee, Nicholas
Cho, Cheol Jun
Black, Alan W
Anumanchipalli, Gopala K.
contents Spoken language models (SLMs) typically discretize speech into high-frame-rate tokens extracted from SSL speech models. As the most successful LMs are based on the Transformer architecture, processing these long token streams with self-attention is expensive, as attention scales quadratically with sequence length. A recent SSL work introduces acoustic tokenization of speech at the syllable level, which is more interpretable and potentially more scalable with significant compression in token lengths (4-5 Hz). Yet, their value for spoken language modeling is not yet fully explored. We present the first systematic study of syllabic tokenization for spoken language modeling, evaluating models on a suite of SLU benchmarks while varying training data scale. Syllabic tokens can match or surpass the previous high-frame rate tokens while significantly cutting training and inference costs, achieving more than a 2x reduction in training time and a 5x reduction in FLOPs. Our findings highlight syllable-level language modeling as a promising path to efficient long-context spoken language models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_26634
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scaling Spoken Language Models with Syllabic Speech Tokenization
Lee, Nicholas
Cho, Cheol Jun
Black, Alan W
Anumanchipalli, Gopala K.
Computation and Language
Audio and Speech Processing
Spoken language models (SLMs) typically discretize speech into high-frame-rate tokens extracted from SSL speech models. As the most successful LMs are based on the Transformer architecture, processing these long token streams with self-attention is expensive, as attention scales quadratically with sequence length. A recent SSL work introduces acoustic tokenization of speech at the syllable level, which is more interpretable and potentially more scalable with significant compression in token lengths (4-5 Hz). Yet, their value for spoken language modeling is not yet fully explored. We present the first systematic study of syllabic tokenization for spoken language modeling, evaluating models on a suite of SLU benchmarks while varying training data scale. Syllabic tokens can match or surpass the previous high-frame rate tokens while significantly cutting training and inference costs, achieving more than a 2x reduction in training time and a 5x reduction in FLOPs. Our findings highlight syllable-level language modeling as a promising path to efficient long-context spoken language models.
title Scaling Spoken Language Models with Syllabic Speech Tokenization
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2509.26634