Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
Fuente:
arXiv
Saved in:
| Main Authors: | Batsuren, Khuyagbaatar, Vylomova, Ekaterina, Dankers, Verna, Delgerbaatar, Tsetsuukhei, Uzan, Omri, Pinter, Yuval, Bella, Gábor |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
by: Uzan, Omri, et al.
Published: (2025)
by: Uzan, Omri, et al.
Published: (2025)
Token Alignment via Character Matching for Subword Completion
by: Athiwaratkun, Ben, et al.
Published: (2024)
by: Athiwaratkun, Ben, et al.
Published: (2024)
Understanding Subword Compositionality of Large Language Models
by: Peng, Qiwei, et al.
Published: (2025)
by: Peng, Qiwei, et al.
Published: (2025)
Team Ryu's Submission to SIGMORPHON 2024 Shared Task on Subword Tokenization
by: Li, Zilong
Published: (2024)
by: Li, Zilong
Published: (2024)
Tokenization Is More Than Compression
by: Schmidt, Craig W., et al.
Published: (2024)
by: Schmidt, Craig W., et al.
Published: (2024)
Greed is All You Need: An Evaluation of Tokenizer Inference Methods
by: Uzan, Omri, et al.
Published: (2024)
by: Uzan, Omri, et al.
Published: (2024)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
by: Deiseroth, Björn, et al.
Published: (2024)
by: Deiseroth, Björn, et al.
Published: (2024)
Optimal Turkish Subword Strategies at Scale: Systematic Evaluation of Data, Vocabulary, Morphology Interplay
by: Altinok, Duygu
Published: (2026)
by: Altinok, Duygu
Published: (2026)
LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation
by: Teklehaymanot, Hailay, et al.
Published: (2026)
by: Teklehaymanot, Hailay, et al.
Published: (2026)
Generating bilingual example sentences with large language models as lexicography assistants
by: Merx, Raphael, et al.
Published: (2024)
by: Merx, Raphael, et al.
Published: (2024)
Subwords as Skills: Tokenization for Sparse-Reward Reinforcement Learning
by: Yunis, David, et al.
Published: (2023)
by: Yunis, David, et al.
Published: (2023)
TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar
by: Li, Yinxi, et al.
Published: (2025)
by: Li, Yinxi, et al.
Published: (2025)
State-of-the-art generalisation research in NLP: A taxonomy and review
by: Hupkes, Dieuwke, et al.
Published: (2022)
by: Hupkes, Dieuwke, et al.
Published: (2022)
Subword Tokenization Strategies for Kurdish Word Embeddings
by: Salehi, Ali, et al.
Published: (2025)
by: Salehi, Ali, et al.
Published: (2025)
Assessing the Importance of Frequency versus Compositionality for Subword-based Tokenization in NMT
by: Wolleb, Benoist, et al.
Published: (2023)
by: Wolleb, Benoist, et al.
Published: (2023)
Tokenization Disparities as Infrastructure Bias: How Subword Systems Create Inequities in LLM Access and Efficiency
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025)
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025)
OpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource Languages
by: Merx, Raphaël, et al.
Published: (2025)
by: Merx, Raphaël, et al.
Published: (2025)
MoVoC: Morphology-Aware Subword Construction for Geez Script Languages
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025)
by: Teklehaymanot, Hailay Kidu, et al.
Published: (2025)
SampoNLP: A Self-Referential Toolkit for Morphological Analysis of Subword Tokenizers
by: Chelombitko, Iaroslav, et al.
Published: (2026)
by: Chelombitko, Iaroslav, et al.
Published: (2026)
Subword-Based Comparative Linguistics across 242 Languages Using Wikipedia Glottosets
by: Chelombitko, Iaroslav, et al.
Published: (2026)
by: Chelombitko, Iaroslav, et al.
Published: (2026)
Low-resource Machine Translation: what for? who for? An observational study on a dedicated Tetun language translation service
by: Merx, Raphael, et al.
Published: (2024)
by: Merx, Raphael, et al.
Published: (2024)
Evaluating Subword Tokenization Techniques for Bengali: A Benchmark Study with BengaliBPE
by: Patwary, Firoj Ahmmed, et al.
Published: (2025)
by: Patwary, Firoj Ahmmed, et al.
Published: (2025)
Tokenization Falling Short: On Subword Robustness in Large Language Models
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
Distributional Properties of Subword Regularization
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Lexically Grounded Subword Segmentation
by: Libovický, Jindřich, et al.
Published: (2024)
by: Libovický, Jindřich, et al.
Published: (2024)
Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder
by: Dai, Yusheng, et al.
Published: (2023)
by: Dai, Yusheng, et al.
Published: (2023)
Evaluating Morphological Plausibility of Subword Tokenization via Statistical Alignment with Morpho-Syntactic Features
by: Stephen, Abishek, et al.
Published: (2026)
by: Stephen, Abishek, et al.
Published: (2026)
Subword Embedding from Bytes Gains Privacy without Sacrificing Accuracy and Complexity
by: Zhang, Mengjiao, et al.
Published: (2024)
by: Zhang, Mengjiao, et al.
Published: (2024)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
by: Ovalle, Anaelia, et al.
Published: (2023)
by: Ovalle, Anaelia, et al.
Published: (2023)
Express Your Doubts -- Probabilistic World Modeling Should not be Based on Token logprobs
by: Wagner, Eitan, et al.
Published: (2025)
by: Wagner, Eitan, et al.
Published: (2025)
ByteSpan: Information-Driven Subword Tokenisation
by: Goriely, Zébulon, et al.
Published: (2025)
by: Goriely, Zébulon, et al.
Published: (2025)
Generalisation First, Memorisation Second? Memorisation Localisation for Natural Language Classification Tasks
by: Dankers, Verna, et al.
Published: (2024)
by: Dankers, Verna, et al.
Published: (2024)
Memorization Inheritance in Sequence-Level Knowledge Distillation for Neural Machine Translation
by: Dankers, Verna, et al.
Published: (2025)
by: Dankers, Verna, et al.
Published: (2025)
Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation
by: Gigant, Théo, et al.
Published: (2026)
by: Gigant, Théo, et al.
Published: (2026)
Existential Definability over the Subword Ordering
by: Baumann, Pascal, et al.
Published: (2022)
by: Baumann, Pascal, et al.
Published: (2022)
Learning Mutually Informed Representations for Characters and Subwords
by: Wang, Yilin, et al.
Published: (2023)
by: Wang, Yilin, et al.
Published: (2023)
Linguistic Laws Meet Protein Sequences: A Comparative Analysis of Subword Tokenization Methods
by: Suyunu, Burak, et al.
Published: (2024)
by: Suyunu, Burak, et al.
Published: (2024)
Subword models struggle with word learning, but surprisal hides it
by: Bunzeck, Bastian, et al.
Published: (2025)
by: Bunzeck, Bastian, et al.
Published: (2025)
The Learning Dynamics of Subword Segmentation for Morphologically Diverse Languages
by: Meyer, Francois, et al.
Published: (2025)
by: Meyer, Francois, et al.
Published: (2025)
Morphological Typology in BPE Subword Productivity and Language Modeling
by: Parra, Iñigo
Published: (2024)
by: Parra, Iñigo
Published: (2024)
Similar Items
-
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
by: Uzan, Omri, et al.
Published: (2025) -
Token Alignment via Character Matching for Subword Completion
by: Athiwaratkun, Ben, et al.
Published: (2024) -
Understanding Subword Compositionality of Large Language Models
by: Peng, Qiwei, et al.
Published: (2025) -
Team Ryu's Submission to SIGMORPHON 2024 Shared Task on Subword Tokenization
by: Li, Zilong
Published: (2024) -
Tokenization Is More Than Compression
by: Schmidt, Craig W., et al.
Published: (2024)