The Foundations of Tokenization: Statistical and Computational Concerns
Fuente:
arXiv
Saved in:
| Main Authors: | Gastaldi, Juan Luis, Terilla, John, Malagutti, Luca, DuSell, Brian, Vieira, Tim, Cotterell, Ryan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Language Models over Tokens to Language Models over Characters
by: Vieira, Tim, et al.
Published: (2024)
by: Vieira, Tim, et al.
Published: (2024)
On the Proper Treatment of Tokenization in Psycholinguistics
by: Giulianelli, Mario, et al.
Published: (2024)
by: Giulianelli, Mario, et al.
Published: (2024)
Bearing Syntactic Fruit with Stack-Augmented Neural Networks
by: DuSell, Brian, et al.
Published: (2025)
by: DuSell, Brian, et al.
Published: (2025)
Better Estimation of the Kullback--Leibler Divergence Between Language Models
by: Amini, Afra, et al.
Published: (2025)
by: Amini, Afra, et al.
Published: (2025)
Direct Preference Optimization with an Offset
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
Training Neural Networks as Recognizers of Formal Languages
by: Butoi, Alexandra, et al.
Published: (2024)
by: Butoi, Alexandra, et al.
Published: (2024)
Language Models over Canonical Byte-Pair Encodings
by: Vieira, Tim, et al.
Published: (2025)
by: Vieira, Tim, et al.
Published: (2025)
Syntactic Control of Language Models by Posterior Inference
by: Xefteri, Vicky, et al.
Published: (2025)
by: Xefteri, Vicky, et al.
Published: (2025)
Variational Best-of-N Alignment
by: Amini, Afra, et al.
Published: (2024)
by: Amini, Afra, et al.
Published: (2024)
Algorithms for Weighted Pushdown Automata
by: Butoi, Alexandra, et al.
Published: (2022)
by: Butoi, Alexandra, et al.
Published: (2022)
Gumbel Counterfactual Generation From Language Models
by: Ravfogel, Shauli, et al.
Published: (2024)
by: Ravfogel, Shauli, et al.
Published: (2024)
Transformers Can Represent $n$-gram Language Models
by: Svete, Anej, et al.
Published: (2024)
by: Svete, Anej, et al.
Published: (2024)
Ensembling Language Models with Sequential Monte Carlo
by: Chan, Robin Shing Moon, et al.
Published: (2026)
by: Chan, Robin Shing Moon, et al.
Published: (2026)
MIO: A Foundation Model on Multimodal Tokens
by: Wang, Zekun, et al.
Published: (2024)
by: Wang, Zekun, et al.
Published: (2024)
Pointwise Mutual Information as a Performance Gauge for Retrieval-Augmented Generation
by: Liu, Tianyu, et al.
Published: (2024)
by: Liu, Tianyu, et al.
Published: (2024)
Reverse-Engineering the Reader
by: Kiegeland, Samuel, et al.
Published: (2024)
by: Kiegeland, Samuel, et al.
Published: (2024)
Fast Controlled Generation from Language Models with Adaptive Weighted Rejection Sampling
by: Lipkin, Benjamin, et al.
Published: (2025)
by: Lipkin, Benjamin, et al.
Published: (2025)
Learning to Reason Efficiently with A* Post-Training
by: Opedal, Andreas, et al.
Published: (2026)
by: Opedal, Andreas, et al.
Published: (2026)
Fine-Tuning and Evaluating Open-Source Large Language Models for the Army Domain
by: Ruiz, Daniel C., et al.
Published: (2024)
by: Ruiz, Daniel C., et al.
Published: (2024)
Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo
by: Loula, João, et al.
Published: (2025)
by: Loula, João, et al.
Published: (2025)
Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns
by: DuSell, Brian, et al.
Published: (2023)
by: DuSell, Brian, et al.
Published: (2023)
Latency and Token-Aware Test-Time Compute
by: Huang, Jenny Y., et al.
Published: (2025)
by: Huang, Jenny Y., et al.
Published: (2025)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
by: Svete, Anej, et al.
Published: (2026)
by: Svete, Anej, et al.
Published: (2026)
On the Bias of Next-Token Predictors Toward Systematically Inefficient Reasoning: A Shortest-Path Case Study
by: Alberghi, Riccardo, et al.
Published: (2025)
by: Alberghi, Riccardo, et al.
Published: (2025)
Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
Revisiting Graph-Tokenizing Large Language Models: A Systematic Evaluation of Graph Token Understanding
by: Zhang, Zhongjian, et al.
Published: (2026)
by: Zhang, Zhongjian, et al.
Published: (2026)
Lossless Token Sequence Compression via Meta-Tokens
by: Harvill, John, et al.
Published: (2025)
by: Harvill, John, et al.
Published: (2025)
Efficient Joint Prediction of Multiple Future Tokens
by: Ahn, Kwangjun, et al.
Published: (2025)
by: Ahn, Kwangjun, et al.
Published: (2025)
Scalable Token-Level Hallucination Detection in Large Language Models
by: Min, Rui, et al.
Published: (2026)
by: Min, Rui, et al.
Published: (2026)
Hierarchical Multi-Label Classification of Online Vaccine Concerns
by: Zhu, Chloe Qinyu, et al.
Published: (2024)
by: Zhu, Chloe Qinyu, et al.
Published: (2024)
TokenButler: Token Importance is Predictable
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
Learning to Route LLMs with Confidence Tokens
by: Chuang, Yu-Neng, et al.
Published: (2024)
by: Chuang, Yu-Neng, et al.
Published: (2024)
Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods
by: Hu, Xinyang, et al.
Published: (2024)
by: Hu, Xinyang, et al.
Published: (2024)
Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
by: Farnik, Lucy, et al.
Published: (2025)
by: Farnik, Lucy, et al.
Published: (2025)
Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
by: Zhao, Minda, et al.
Published: (2026)
by: Zhao, Minda, et al.
Published: (2026)
Adversarial Tokenization
by: Geh, Renato Lui, et al.
Published: (2025)
by: Geh, Renato Lui, et al.
Published: (2025)
On Efficient and Statistical Quality Estimation for Data Annotation
by: Klie, Jan-Christoph, et al.
Published: (2024)
by: Klie, Jan-Christoph, et al.
Published: (2024)
RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably
by: Du, Yufeng, et al.
Published: (2026)
by: Du, Yufeng, et al.
Published: (2026)
Rep2Text: Decoding Full Text from a Single LLM Token Representation
by: Zhao, Haiyan, et al.
Published: (2025)
by: Zhao, Haiyan, et al.
Published: (2025)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
by: Ovalle, Anaelia, et al.
Published: (2023)
by: Ovalle, Anaelia, et al.
Published: (2023)
Similar Items
-
From Language Models over Tokens to Language Models over Characters
by: Vieira, Tim, et al.
Published: (2024) -
On the Proper Treatment of Tokenization in Psycholinguistics
by: Giulianelli, Mario, et al.
Published: (2024) -
Bearing Syntactic Fruit with Stack-Augmented Neural Networks
by: DuSell, Brian, et al.
Published: (2025) -
Better Estimation of the Kullback--Leibler Divergence Between Language Models
by: Amini, Afra, et al.
Published: (2025) -
Direct Preference Optimization with an Offset
by: Amini, Afra, et al.
Published: (2024)