Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Alqahtani, Sawsan, Nayeem, Mir Tafseer, Laskar, Md Tahmid Rahman, Mohiuddin, Tasnim, Bari, M Saiful
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908784877633536
author Alqahtani, Sawsan
Nayeem, Mir Tafseer
Laskar, Md Tahmid Rahman
Mohiuddin, Tasnim
Bari, M Saiful
author_facet Alqahtani, Sawsan
Nayeem, Mir Tafseer
Laskar, Md Tahmid Rahman
Mohiuddin, Tasnim
Bari, M Saiful
contents Tokenization underlies every large language model, yet it remains an under-theorized and inconsistently designed component. Common subword approaches such as Byte Pair Encoding (BPE) offer scalability but often misalign with linguistic structure, amplify bias, and waste capacity across languages and domains. This paper reframes tokenization as a core modeling decision rather than a preprocessing step. We argue for a context-aware framework that integrates tokenizer and model co-design, guided by linguistic, domain, and deployment considerations. Standardized evaluation and transparent reporting are essential to make tokenization choices accountable and comparable. Treating tokenization as a core design problem, not a technical afterthought, can yield language technologies that are fairer, more efficient, and more adaptable.
format Preprint
id arxiv_https___arxiv_org_abs_2601_13260
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models
Alqahtani, Sawsan
Nayeem, Mir Tafseer
Laskar, Md Tahmid Rahman
Mohiuddin, Tasnim
Bari, M Saiful
Computation and Language
Artificial Intelligence
Machine Learning
Tokenization underlies every large language model, yet it remains an under-theorized and inconsistently designed component. Common subword approaches such as Byte Pair Encoding (BPE) offer scalability but often misalign with linguistic structure, amplify bias, and waste capacity across languages and domains. This paper reframes tokenization as a core modeling decision rather than a preprocessing step. We argue for a context-aware framework that integrates tokenizer and model co-design, guided by linguistic, domain, and deployment considerations. Standardized evaluation and transparent reporting are essential to make tokenization choices accountable and comparable. Treating tokenization as a core design problem, not a technical afterthought, can yield language technologies that are fairer, more efficient, and more adaptable.
title Stop Taking Tokenizers for Granted: They Are Core Design Decisions in Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.13260