ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Visser, Nicol, Malan, Simon, Slabbert, Danel, Kamper, Herman
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917277938483200
author Visser, Nicol
Malan, Simon
Slabbert, Danel
Kamper, Herman
author_facet Visser, Nicol
Malan, Simon
Slabbert, Danel
Kamper, Herman
contents Pure speech language models aim to learn language directly from raw audio without textual resources. A key challenge is that discrete tokens from self-supervised speech encoders result in excessively long sequences, motivating recent work on syllable-like units. However, methods like Sylber and SyllableLM rely on intricate multi-stage training pipelines. We propose ZeroSyl, a simple training-free method to extract syllable boundaries and embeddings directly from a frozen WavLM model. Using L2 norms of features in WavLM's intermediate layers, ZeroSyl achieves competitive syllable segmentation performance. The resulting segments are mean-pooled, discretized using K-means, and used to train a language model. ZeroSyl outperforms prior syllabic tokenizers across lexical, syntactic, and narrative benchmarks. Scaling experiments show that while finer-grained units are beneficial for lexical tasks, our discovered syllabic units exhibit better scaling behavior for syntactic modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15537
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
Visser, Nicol
Malan, Simon
Slabbert, Danel
Kamper, Herman
Computation and Language
Audio and Speech Processing
Pure speech language models aim to learn language directly from raw audio without textual resources. A key challenge is that discrete tokens from self-supervised speech encoders result in excessively long sequences, motivating recent work on syllable-like units. However, methods like Sylber and SyllableLM rely on intricate multi-stage training pipelines. We propose ZeroSyl, a simple training-free method to extract syllable boundaries and embeddings directly from a frozen WavLM model. Using L2 norms of features in WavLM's intermediate layers, ZeroSyl achieves competitive syllable segmentation performance. The resulting segments are mean-pooled, discretized using K-means, and used to train a language model. ZeroSyl outperforms prior syllabic tokenizers across lexical, syntactic, and narrative benchmarks. Scaling experiments show that while finer-grained units are beneficial for lexical tasks, our discovered syllabic units exhibit better scaling behavior for syntactic modeling.
title ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
topic Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2602.15537