The Foundations of Tokenization: Statistical and Computational Concerns

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gastaldi, Juan Luis, Terilla, John, Malagutti, Luca, DuSell, Brian, Vieira, Tim, Cotterell, Ryan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912306121670656
author Gastaldi, Juan Luis
Terilla, John
Malagutti, Luca
DuSell, Brian
Vieira, Tim
Cotterell, Ryan
author_facet Gastaldi, Juan Luis
Terilla, John
Malagutti, Luca
DuSell, Brian
Vieira, Tim
Cotterell, Ryan
contents Tokenization - the practice of converting strings of characters from an alphabet into sequences of tokens over a vocabulary - is a critical step in the NLP pipeline. The use of token representations is widely credited with increased model performance but is also the source of many undesirable behaviors, such as spurious ambiguity or inconsistency. Despite its recognized importance as a standard representation method in NLP, the theoretical underpinnings of tokenization are not yet fully understood. In particular, the impact of tokenization on language model estimation has been investigated primarily through empirical means. The present paper contributes to addressing this theoretical gap by proposing a unified formal framework for representing and analyzing tokenizer models. Based on the category of stochastic maps, this framework enables us to establish general conditions for a principled use of tokenizers and, most importantly, the necessary and sufficient conditions for a tokenizer model to preserve the consistency of statistical estimators. In addition, we discuss statistical and computational concerns crucial for designing and implementing tokenizer models, such as inconsistency, ambiguity, finiteness, and sequentiality. The framework and results advanced in this paper contribute to building robust theoretical foundations for representations in neural language modeling that can inform future theoretical and empirical research.
format Preprint
id arxiv_https___arxiv_org_abs_2407_11606
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Foundations of Tokenization: Statistical and Computational Concerns
Gastaldi, Juan Luis
Terilla, John
Malagutti, Luca
DuSell, Brian
Vieira, Tim
Cotterell, Ryan
Computation and Language
Artificial Intelligence
Machine Learning
Tokenization - the practice of converting strings of characters from an alphabet into sequences of tokens over a vocabulary - is a critical step in the NLP pipeline. The use of token representations is widely credited with increased model performance but is also the source of many undesirable behaviors, such as spurious ambiguity or inconsistency. Despite its recognized importance as a standard representation method in NLP, the theoretical underpinnings of tokenization are not yet fully understood. In particular, the impact of tokenization on language model estimation has been investigated primarily through empirical means. The present paper contributes to addressing this theoretical gap by proposing a unified formal framework for representing and analyzing tokenizer models. Based on the category of stochastic maps, this framework enables us to establish general conditions for a principled use of tokenizers and, most importantly, the necessary and sufficient conditions for a tokenizer model to preserve the consistency of statistical estimators. In addition, we discuss statistical and computational concerns crucial for designing and implementing tokenizer models, such as inconsistency, ambiguity, finiteness, and sequentiality. The framework and results advanced in this paper contribute to building robust theoretical foundations for representations in neural language modeling that can inform future theoretical and empirical research.
title The Foundations of Tokenization: Statistical and Computational Concerns
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.11606