Leading Whitespaces of Language Models' Subword Vocabulary Pose a Confound for Calculating Word Probabilities
Fuente:
arXiv
Saved in:
| Main Authors: | Oh, Byung-Doh, Schuler, William |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Impact of Token Granularity on the Predictive Power of Language Model Surprisal
by: Oh, Byung-Doh, et al.
Published: (2024)
by: Oh, Byung-Doh, et al.
Published: (2024)
How Well Does First-Token Entropy Approximate Word Entropy as a Psycholinguistic Predictor?
by: Clark, Christian, et al.
Published: (2025)
by: Clark, Christian, et al.
Published: (2025)
The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage
by: Oh, Byung-Doh, et al.
Published: (2025)
by: Oh, Byung-Doh, et al.
Published: (2025)
Frequency Explains the Inverse Correlation of Large Language Models' Size, Training Data Amount, and Surprisal's Fit to Reading Times
by: Oh, Byung-Doh, et al.
Published: (2024)
by: Oh, Byung-Doh, et al.
Published: (2024)
Linear Recency Bias During Training Improves Transformers' Fit to Reading Times
by: Clark, Christian, et al.
Published: (2024)
by: Clark, Christian, et al.
Published: (2024)
Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal
by: Nair, Sathvik, et al.
Published: (2026)
by: Nair, Sathvik, et al.
Published: (2026)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
by: Zouhar, Vilém
Published: (2024)
by: Zouhar, Vilém
Published: (2024)
Subword Tokenization Strategies for Kurdish Word Embeddings
by: Salehi, Ali, et al.
Published: (2025)
by: Salehi, Ali, et al.
Published: (2025)
Morphological Typology in BPE Subword Productivity and Language Modeling
by: Parra, Iñigo
Published: (2024)
by: Parra, Iñigo
Published: (2024)
Tokenization Falling Short: On Subword Robustness in Large Language Models
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
Optimal Turkish Subword Strategies at Scale: Systematic Evaluation of Data, Vocabulary, Morphology Interplay
by: Altinok, Duygu
Published: (2026)
by: Altinok, Duygu
Published: (2026)
so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMs
by: Bhyravajjula, Sriharsh, et al.
Published: (2025)
by: Bhyravajjula, Sriharsh, et al.
Published: (2025)
Understanding Subword Compositionality of Large Language Models
by: Peng, Qiwei, et al.
Published: (2025)
by: Peng, Qiwei, et al.
Published: (2025)
The Learning Dynamics of Subword Segmentation for Morphologically Diverse Languages
by: Meyer, Francois, et al.
Published: (2025)
by: Meyer, Francois, et al.
Published: (2025)
On the Effect of (Near) Duplicate Subwords in Language Modelling
by: Schäfer, Anton, et al.
Published: (2024)
by: Schäfer, Anton, et al.
Published: (2024)
Surprisal from Larger Transformer-based Language Models Predicts fMRI Data More Poorly
by: Lin, Yi-Chien, et al.
Published: (2025)
by: Lin, Yi-Chien, et al.
Published: (2025)
How to Compute the Probability of a Word
by: Pimentel, Tiago, et al.
Published: (2024)
by: Pimentel, Tiago, et al.
Published: (2024)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
by: Batsuren, Khuyagbaatar, et al.
Published: (2024)
by: Batsuren, Khuyagbaatar, et al.
Published: (2024)
Distributional Properties of Subword Regularization
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Lexically Grounded Subword Segmentation
by: Libovický, Jindřich, et al.
Published: (2024)
by: Libovický, Jindřich, et al.
Published: (2024)
Enhancing Sindhi Word Segmentation using Subword Representation Learning and Position-aware Self-attention
by: Ali, Wazir, et al.
Published: (2020)
by: Ali, Wazir, et al.
Published: (2020)
Using Perspectival Words Is Harder Than Vocabulary Words for Humans and Even More So for Multimodal Language Models
by: Dong, Dota Tianai, et al.
Published: (2025)
by: Dong, Dota Tianai, et al.
Published: (2025)
Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation
by: Gigant, Théo, et al.
Published: (2026)
by: Gigant, Théo, et al.
Published: (2026)
Can Pretrained Language Models Derive Correct Semantics from Corrupt Subwords under Noise?
by: Li, Xinzhe, et al.
Published: (2023)
by: Li, Xinzhe, et al.
Published: (2023)
Position Paper On Diagnostic Uncertainty Estimation from Large Language Models: Next-Word Probability Is Not Pre-test Probability
by: Gao, Yanjun, et al.
Published: (2024)
by: Gao, Yanjun, et al.
Published: (2024)
Mitigating Hidden Confounding by Progressive Confounder Imputation via Large Language Models
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
Tomato, Tomahto, Tomate: Do Multilingual Language Models Understand Based on Subword-Level Semantic Concepts?
by: Zhang, Crystina, et al.
Published: (2024)
by: Zhang, Crystina, et al.
Published: (2024)
ByteSpan: Information-Driven Subword Tokenisation
by: Goriely, Zébulon, et al.
Published: (2025)
by: Goriely, Zébulon, et al.
Published: (2025)
The Frequency Confound in Language-Model Surprisal and Metaphor Novelty
by: Momen, Omar, et al.
Published: (2026)
by: Momen, Omar, et al.
Published: (2026)
Vectors from Larger Language Models Predict Human Reading Time and fMRI Data More Poorly when Dimensionality Expansion is Controlled
by: Lin, Yi-Chien, et al.
Published: (2025)
by: Lin, Yi-Chien, et al.
Published: (2025)
Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning
by: Kim, Nayeon, et al.
Published: (2025)
by: Kim, Nayeon, et al.
Published: (2025)
Pretraining Language Models with Subword Regularization: An Empirical Study of BPE Dropout in Low-Resource NLP
by: Visser, Ruan, et al.
Published: (2026)
by: Visser, Ruan, et al.
Published: (2026)
World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language Models
by: Ma, Ziqiao, et al.
Published: (2023)
by: Ma, Ziqiao, et al.
Published: (2023)
Subword models struggle with word learning, but surprisal hides it
by: Bunzeck, Bastian, et al.
Published: (2025)
by: Bunzeck, Bastian, et al.
Published: (2025)
Transferring Extreme Subword Style Using Ngram Model-Based Logit Scaling
by: Messner, Craig, et al.
Published: (2025)
by: Messner, Craig, et al.
Published: (2025)
Two-step Automated Cybercrime Coded Word Detection using Multi-level Representation Learning
by: Kim, Yongyeon, et al.
Published: (2024)
by: Kim, Yongyeon, et al.
Published: (2024)
LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation
by: Teklehaymanot, Hailay, et al.
Published: (2026)
by: Teklehaymanot, Hailay, et al.
Published: (2026)
Learning Mutually Informed Representations for Characters and Subwords
by: Wang, Yilin, et al.
Published: (2023)
by: Wang, Yilin, et al.
Published: (2023)
StochasTok: Improving Fine-Grained Subword Understanding in LLMs
by: Sims, Anya, et al.
Published: (2025)
by: Sims, Anya, et al.
Published: (2025)
Beyond Shared Vocabulary: Increasing Representational Word Similarities across Languages for Multilingual Machine Translation
by: Wu, Di, et al.
Published: (2023)
by: Wu, Di, et al.
Published: (2023)
Similar Items
-
The Impact of Token Granularity on the Predictive Power of Language Model Surprisal
by: Oh, Byung-Doh, et al.
Published: (2024) -
How Well Does First-Token Entropy Approximate Word Entropy as a Psycholinguistic Predictor?
by: Clark, Christian, et al.
Published: (2025) -
The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage
by: Oh, Byung-Doh, et al.
Published: (2025) -
Frequency Explains the Inverse Correlation of Large Language Models' Size, Training Data Amount, and Surprisal's Fit to Reading Times
by: Oh, Byung-Doh, et al.
Published: (2024) -
Linear Recency Bias During Training Improves Transformers' Fit to Reading Times
by: Clark, Christian, et al.
Published: (2024)