SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chelombitko, Iaroslav, Chelombitko, Ekaterina, Komissarov, Aleksey
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912809666740224
author Chelombitko, Iaroslav
Chelombitko, Ekaterina
Komissarov, Aleksey
author_facet Chelombitko, Iaroslav
Chelombitko, Ekaterina
Komissarov, Aleksey
contents The quality of subword tokenization is critical for Large Language Models, yet evaluating tokenizers for morphologically rich Uralic languages is hampered by the lack of clean morpheme lexicons. We introduce SampoNLP, a corpus-free toolkit for morphological lexicon creation using MDL-inspired Self-Referential Atomicity Scoring, which filters composite forms through internal structural cues - suited for low-resource settings. Using the high-purity lexicons generated by SampoNLP for Finnish, Hungarian, and Estonian, we conduct a systematic evaluation of BPE tokenizers across a range of vocabulary sizes (8k-256k). We propose a unified metric, the Integrated Performance Score (IPS), to navigate the trade-off between morpheme coverage and over-splitting. By analyzing the IPS curves, we identify the "elbow points" of diminishing returns and provide the first empirically grounded recommendations for optimal vocabulary sizes (k) in these languages. Our study not only offers practical guidance but also quantitatively demonstrates the limitations of standard BPE for highly agglutinative languages. The SampoNLP library and all generated resources are made publicly available: https://github.com/AragonerUA/SampoNLP
format Preprint
id arxiv_https___arxiv_org_abs_2601_04469
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers
Chelombitko, Iaroslav
Chelombitko, Ekaterina
Komissarov, Aleksey
Computation and Language
Information Retrieval
Machine Learning
I.2.7; I.2.6; H.3.1
The quality of subword tokenization is critical for Large Language Models, yet evaluating tokenizers for morphologically rich Uralic languages is hampered by the lack of clean morpheme lexicons. We introduce SampoNLP, a corpus-free toolkit for morphological lexicon creation using MDL-inspired Self-Referential Atomicity Scoring, which filters composite forms through internal structural cues - suited for low-resource settings. Using the high-purity lexicons generated by SampoNLP for Finnish, Hungarian, and Estonian, we conduct a systematic evaluation of BPE tokenizers across a range of vocabulary sizes (8k-256k). We propose a unified metric, the Integrated Performance Score (IPS), to navigate the trade-off between morpheme coverage and over-splitting. By analyzing the IPS curves, we identify the "elbow points" of diminishing returns and provide the first empirically grounded recommendations for optimal vocabulary sizes (k) in these languages. Our study not only offers practical guidance but also quantitatively demonstrates the limitations of standard BPE for highly agglutinative languages. The SampoNLP library and all generated resources are made publicly available: https://github.com/AragonerUA/SampoNLP
title SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers
topic Computation and Language
Information Retrieval
Machine Learning
I.2.7; I.2.6; H.3.1
url https://arxiv.org/abs/2601.04469