StochasTok: Improving Fine-Grained Subword Understanding in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sims, Anya, Foster, Thom, Kaleb, Klara, Nguyen, Tuan-Duy H., Lee, Joseph, Foerster, Jakob N., Teh, Yee Whye, Lu, Cong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915946044588032
author Sims, Anya
Foster, Thom
Kaleb, Klara
Nguyen, Tuan-Duy H.
Lee, Joseph
Foerster, Jakob N.
Teh, Yee Whye
Lu, Cong
author_facet Sims, Anya
Foster, Thom
Kaleb, Klara
Nguyen, Tuan-Duy H.
Lee, Joseph
Foerster, Jakob N.
Teh, Yee Whye
Lu, Cong
contents Subword-level understanding is integral to numerous tasks, including understanding multi-digit numbers, spelling mistakes, abbreviations, rhyming, and wordplay. Despite this, current large language models (LLMs) still struggle disproportionally with simple subword-level tasks like 'How many r's in strawberry?'. A key factor behind these failures is tokenization, which obscures the fine-grained structure of words. Current alternatives, such as character-level and dropout tokenization methods, significantly increase computational costs and provide inconsistent improvements. In this paper we revisit tokenization and introduce StochasTok, a simple, efficient stochastic tokenization scheme that randomly splits tokens during training, allowing LLMs to 'see' their internal structure. Our experiments show that pretraining with StochasTok substantially improves LLMs' downstream performance across multiple subword-level language games, including character counting, substring identification, and math tasks. Furthermore, StochasTok's simplicity allows seamless integration at any stage of the training pipeline; and we demonstrate that post-training with StochasTok can instill improved subword understanding into existing pretrained models, thus avoiding costly pretraining from scratch. These dramatic improvements achieved with a minimal change suggest StochasTok holds exciting potential when applied to larger, more capable models. Code open-sourced at: github.com/anyasims/stochastok.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01687
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle StochasTok: Improving Fine-Grained Subword Understanding in LLMs
Sims, Anya
Foster, Thom
Kaleb, Klara
Nguyen, Tuan-Duy H.
Lee, Joseph
Foerster, Jakob N.
Teh, Yee Whye
Lu, Cong
Computation and Language
Subword-level understanding is integral to numerous tasks, including understanding multi-digit numbers, spelling mistakes, abbreviations, rhyming, and wordplay. Despite this, current large language models (LLMs) still struggle disproportionally with simple subword-level tasks like 'How many r's in strawberry?'. A key factor behind these failures is tokenization, which obscures the fine-grained structure of words. Current alternatives, such as character-level and dropout tokenization methods, significantly increase computational costs and provide inconsistent improvements. In this paper we revisit tokenization and introduce StochasTok, a simple, efficient stochastic tokenization scheme that randomly splits tokens during training, allowing LLMs to 'see' their internal structure. Our experiments show that pretraining with StochasTok substantially improves LLMs' downstream performance across multiple subword-level language games, including character counting, substring identification, and math tasks. Furthermore, StochasTok's simplicity allows seamless integration at any stage of the training pipeline; and we demonstrate that post-training with StochasTok can instill improved subword understanding into existing pretrained models, thus avoiding costly pretraining from scratch. These dramatic improvements achieved with a minimal change suggest StochasTok holds exciting potential when applied to larger, more capable models. Code open-sourced at: github.com/anyasims/stochastok.
title StochasTok: Improving Fine-Grained Subword Understanding in LLMs
topic Computation and Language
url https://arxiv.org/abs/2506.01687