BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization
Fuente:
arXiv
Saved in:
| Main Authors: | Land, Sander, Arnett, Catherine |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
by: Chizhov, Pavel, et al.
Published: (2024)
by: Chizhov, Pavel, et al.
Published: (2024)
BlockBPE: Parallel BPE Tokenization
by: You, Amos
Published: (2025)
by: You, Amos
Published: (2025)
Peek2: Regex-free Byte-level Byte-Pair Encoding Pretokenizer for LLM Inference on Edge Devices
by: Zai, Liu, et al.
Published: (2026)
by: Zai, Liu, et al.
Published: (2026)
Which Pieces Does Unigram Tokenization Really Need?
by: Land, Sander, et al.
Published: (2025)
by: Land, Sander, et al.
Published: (2025)
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
by: Land, Sander, et al.
Published: (2024)
by: Land, Sander, et al.
Published: (2024)
Auditing LLM Benchmarks with Item Response Theory
by: Land, Sander, et al.
Published: (2026)
by: Land, Sander, et al.
Published: (2026)
Scaffold-BPE: Enhancing Byte Pair Encoding for Large Language Models with Simple and Effective Scaffold Token Removal
by: Lian, Haoran, et al.
Published: (2024)
by: Lian, Haoran, et al.
Published: (2024)
Distilling Multilingual Vision-Language Models: When Smaller Models Stay Multilingual
by: Sriratanawilai, Sukrit, et al.
Published: (2025)
by: Sriratanawilai, Sukrit, et al.
Published: (2025)
Disaggregation Reveals Hidden Training Dynamics: The Case of Agreement Attraction
by: Michaelov, James A., et al.
Published: (2025)
by: Michaelov, James A., et al.
Published: (2025)
Why do language models perform worse for morphologically complex languages?
by: Arnett, Catherine, et al.
Published: (2024)
by: Arnett, Catherine, et al.
Published: (2024)
Batching BPE Tokenization Merges
by: Morgan, Alexander P.
Published: (2024)
by: Morgan, Alexander P.
Published: (2024)
Bit-level BPE: Below the byte boundary
by: Moon, Sangwhan, et al.
Published: (2025)
by: Moon, Sangwhan, et al.
Published: (2025)
Byte BPE Tokenization as an Inverse string Homomorphism
by: Geng, Saibo, et al.
Published: (2024)
by: Geng, Saibo, et al.
Published: (2024)
Constructing a BPE Tokenization DFA
by: Berglund, Martin, et al.
Published: (2024)
by: Berglund, Martin, et al.
Published: (2024)
From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
by: Dolga, Rares, et al.
Published: (2025)
by: Dolga, Rares, et al.
Published: (2025)
AdaptBPE: From General Purpose to Specialized Tokenizers
by: Liyanage, Vijini, et al.
Published: (2026)
by: Liyanage, Vijini, et al.
Published: (2026)
Morphological Typology in BPE Subword Productivity and Language Modeling
by: Parra, Iñigo
Published: (2024)
by: Parra, Iñigo
Published: (2024)
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens
by: Zhao, Zhenyu, et al.
Published: (2026)
by: Zhao, Zhenyu, et al.
Published: (2026)
Multilingual Language Models Encode Script Over Linguistic Structure
by: Verma, Aastha A K, et al.
Published: (2026)
by: Verma, Aastha A K, et al.
Published: (2026)
SuperBPE: Space Travel for Language Models
by: Liu, Alisa, et al.
Published: (2025)
by: Liu, Alisa, et al.
Published: (2025)
Evaluating Morphological Alignment of Tokenizers in 70 Languages
by: Arnett, Catherine, et al.
Published: (2025)
by: Arnett, Catherine, et al.
Published: (2025)
Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models
by: Sawada, Tomohiro, et al.
Published: (2025)
by: Sawada, Tomohiro, et al.
Published: (2025)
LiteToken: Removing Intermediate Merge Residues From BPE Tokenizers
by: Sun, Yike, et al.
Published: (2026)
by: Sun, Yike, et al.
Published: (2026)
Revenge of the Fallen? Recurrent Models Match Transformers at Predicting Human Language Comprehension Metrics
by: Michaelov, James A., et al.
Published: (2024)
by: Michaelov, James A., et al.
Published: (2024)
Weight Tying Biases Token Embeddings Towards the Output Space
by: Lopardo, Antonio, et al.
Published: (2026)
by: Lopardo, Antonio, et al.
Published: (2026)
A Bit of a Problem: Measurement Disparities in Dataset Sizes Across Languages
by: Arnett, Catherine, et al.
Published: (2024)
by: Arnett, Catherine, et al.
Published: (2024)
Evaluating Subword Tokenization Techniques for Bengali: A Benchmark Study with BengaliBPE
by: Patwary, Firoj Ahmmed, et al.
Published: (2025)
by: Patwary, Firoj Ahmmed, et al.
Published: (2025)
Adaptive BPE Tokenization for Enhanced Vocabulary Adaptation in Finetuning Pretrained Language Models
by: Balde, Gunjan, et al.
Published: (2024)
by: Balde, Gunjan, et al.
Published: (2024)
SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models
by: Kim, SungHo, et al.
Published: (2026)
by: Kim, SungHo, et al.
Published: (2026)
Romanization Encoding For Multilingual ASR
by: Ding, Wen, et al.
Published: (2024)
by: Ding, Wen, et al.
Published: (2024)
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
by: Kadamba, Venu Gopal, et al.
Published: (2026)
by: Kadamba, Venu Gopal, et al.
Published: (2026)
Explaining and Mitigating Crosslingual Tokenizer Inequities
by: Arnett, Catherine, et al.
Published: (2025)
by: Arnett, Catherine, et al.
Published: (2025)
Goldfish: Monolingual Language Models for 350 Languages
by: Chang, Tyler A., et al.
Published: (2024)
by: Chang, Tyler A., et al.
Published: (2024)
Different Tokenization Schemes Lead to Comparable Performance in Spanish Number Agreement
by: Arnett, Catherine, et al.
Published: (2024)
by: Arnett, Catherine, et al.
Published: (2024)
Toxicity of the Commons: Curating Open-Source Pre-Training Data
by: Arnett, Catherine, et al.
Published: (2024)
by: Arnett, Catherine, et al.
Published: (2024)
Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
by: Vemula, Saketh Reddy, et al.
Published: (2025)
by: Vemula, Saketh Reddy, et al.
Published: (2025)
Stay Hungry, Stay Foolish: On the Extended Reading Articles Generation with LLMs
by: Liou, Yow-Fu, et al.
Published: (2025)
by: Liou, Yow-Fu, et al.
Published: (2025)
On the Acquisition of Shared Grammatical Representations in Bilingual Language Models
by: Arnett, Catherine, et al.
Published: (2025)
by: Arnett, Catherine, et al.
Published: (2025)
Understanding Likelihood Over-optimisation in Direct Alignment Algorithms
by: Shi, Zhengyan, et al.
Published: (2024)
by: Shi, Zhengyan, et al.
Published: (2024)
Similar Items
-
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
by: Chizhov, Pavel, et al.
Published: (2024) -
BlockBPE: Parallel BPE Tokenization
by: You, Amos
Published: (2025) -
Peek2: Regex-free Byte-level Byte-Pair Encoding Pretokenizer for LLM Inference on Edge Devices
by: Zai, Liu, et al.
Published: (2026) -
Which Pieces Does Unigram Tokenization Really Need?
by: Land, Sander, et al.
Published: (2025) -
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
by: Land, Sander, et al.
Published: (2024)