LiteToken: Removing Intermediate Merge Residues From BPE Tokenizers

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Yike, Yang, Haotong, Lin, Zhouchen, Zhang, Muhan
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917248080281600
author Sun, Yike
Yang, Haotong
Lin, Zhouchen
Zhang, Muhan
author_facet Sun, Yike
Yang, Haotong
Lin, Zhouchen
Zhang, Muhan
contents Tokenization is fundamental to how language models represent and process text, yet the behavior of widely used BPE tokenizers has received far less study than model architectures and training. In this paper, we investigate intermediate merge residues in BPE vocabularies: tokens that are frequent during merge learning so that retained in the final vocabulary, but are mostly further merged and rarely emitted when tokenizing the corpus during tokenizer usage. Such low-frequency tokens not only waste vocabulary capacity but also increase vulnerability to adversarial or atypical inputs. We present a systematic empirical characterization of this phenomenon across commonly used tokenizers and introduce LiteToken, a simple method for removing residue tokens. Because the affected tokens are rarely used, pretrained models can often accommodate the modified tokenizer without additional fine-tuning. Experiments show that LiteToken reduces token fragmentation, reduces parameters, and improves robustness to noisy or misspelled inputs, while preserving overall performance.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04706
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LiteToken: Removing Intermediate Merge Residues From BPE Tokenizers
Sun, Yike
Yang, Haotong
Lin, Zhouchen
Zhang, Muhan
Computation and Language
Tokenization is fundamental to how language models represent and process text, yet the behavior of widely used BPE tokenizers has received far less study than model architectures and training. In this paper, we investigate intermediate merge residues in BPE vocabularies: tokens that are frequent during merge learning so that retained in the final vocabulary, but are mostly further merged and rarely emitted when tokenizing the corpus during tokenizer usage. Such low-frequency tokens not only waste vocabulary capacity but also increase vulnerability to adversarial or atypical inputs. We present a systematic empirical characterization of this phenomenon across commonly used tokenizers and introduce LiteToken, a simple method for removing residue tokens. Because the affected tokens are rarely used, pretrained models can often accommodate the modified tokenizer without additional fine-tuning. Experiments show that LiteToken reduces token fragmentation, reduces parameters, and improves robustness to noisy or misspelled inputs, while preserving overall performance.
title LiteToken: Removing Intermediate Merge Residues From BPE Tokenizers
topic Computation and Language
url https://arxiv.org/abs/2602.04706