From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
Fuente:
arXiv
Saved in:
| Main Authors: | Dolga, Rares, Maystre, Lucas, Berariu, Tudor, Barber, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unifying Linear-Time Attention via Latent Probabilistic Modelling
by: Dolga, Rares, et al.
Published: (2024)
by: Dolga, Rares, et al.
Published: (2024)
Incremental Sequence Classification with Temporal Consistency
by: Maystre, Lucas, et al.
Published: (2025)
by: Maystre, Lucas, et al.
Published: (2025)
When Embedding Models Meet: Procrustes Bounds and Applications
by: Maystre, Lucas, et al.
Published: (2025)
by: Maystre, Lucas, et al.
Published: (2025)
BlockBPE: Parallel BPE Tokenization
by: You, Amos
Published: (2025)
by: You, Amos
Published: (2025)
AdaptBPE: From General Purpose to Specialized Tokenizers
by: Liyanage, Vijini, et al.
Published: (2026)
by: Liyanage, Vijini, et al.
Published: (2026)
LiteToken: Removing Intermediate Merge Residues From BPE Tokenizers
by: Sun, Yike, et al.
Published: (2026)
by: Sun, Yike, et al.
Published: (2026)
Batching BPE Tokenization Merges
by: Morgan, Alexander P.
Published: (2024)
by: Morgan, Alexander P.
Published: (2024)
Constructing a BPE Tokenization DFA
by: Berglund, Martin, et al.
Published: (2024)
by: Berglund, Martin, et al.
Published: (2024)
Byte BPE Tokenization as an Inverse string Homomorphism
by: Geng, Saibo, et al.
Published: (2024)
by: Geng, Saibo, et al.
Published: (2024)
Log Summarisation for Defect Evolution Analysis
by: Dolga, Rares, et al.
Published: (2024)
by: Dolga, Rares, et al.
Published: (2024)
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
by: Chizhov, Pavel, et al.
Published: (2024)
by: Chizhov, Pavel, et al.
Published: (2024)
Evaluating Subword Tokenization Techniques for Bengali: A Benchmark Study with BengaliBPE
by: Patwary, Firoj Ahmmed, et al.
Published: (2025)
by: Patwary, Firoj Ahmmed, et al.
Published: (2025)
Adaptive BPE Tokenization for Enhanced Vocabulary Adaptation in Finetuning Pretrained Language Models
by: Balde, Gunjan, et al.
Published: (2024)
by: Balde, Gunjan, et al.
Published: (2024)
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
by: Kadamba, Venu Gopal, et al.
Published: (2026)
by: Kadamba, Venu Gopal, et al.
Published: (2026)
Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
by: Vemula, Saketh Reddy, et al.
Published: (2025)
by: Vemula, Saketh Reddy, et al.
Published: (2025)
Every Character Counts: From Vulnerability to Defense in Phishing Detection
by: Chiper, Maria, et al.
Published: (2025)
by: Chiper, Maria, et al.
Published: (2025)
Toward Autonomous UI Exploration: The UIExplorer Benchmark
by: Nica, Andrei Cristian, et al.
Published: (2025)
by: Nica, Andrei Cristian, et al.
Published: (2025)
RotRNN: Modelling Long Sequences with Rotations
by: Biegun, Kai, et al.
Published: (2024)
by: Biegun, Kai, et al.
Published: (2024)
Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?
by: Hayase, Jonathan, et al.
Published: (2024)
by: Hayase, Jonathan, et al.
Published: (2024)
Scaffold-BPE: Enhancing Byte Pair Encoding for Large Language Models with Simple and Effective Scaffold Token Removal
by: Lian, Haoran, et al.
Published: (2024)
by: Lian, Haoran, et al.
Published: (2024)
Bit-level BPE: Below the byte boundary
by: Moon, Sangwhan, et al.
Published: (2025)
by: Moon, Sangwhan, et al.
Published: (2025)
Morphological Typology in BPE Subword Productivity and Language Modeling
by: Parra, Iñigo
Published: (2024)
by: Parra, Iñigo
Published: (2024)
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
by: Asgari, Ehsaneddin, et al.
Published: (2025)
by: Asgari, Ehsaneddin, et al.
Published: (2025)
Spelling-out is not Straightforward: LLMs' Capability of Tokenization from Token to Characters
by: Hiraoka, Tatsuya, et al.
Published: (2025)
by: Hiraoka, Tatsuya, et al.
Published: (2025)
SuperBPE: Space Travel for Language Models
by: Liu, Alisa, et al.
Published: (2025)
by: Liu, Alisa, et al.
Published: (2025)
BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization
by: Land, Sander, et al.
Published: (2025)
by: Land, Sander, et al.
Published: (2025)
From Language Models over Tokens to Language Models over Characters
by: Vieira, Tim, et al.
Published: (2024)
by: Vieira, Tim, et al.
Published: (2024)
Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models
by: Sawada, Tomohiro, et al.
Published: (2025)
by: Sawada, Tomohiro, et al.
Published: (2025)
Adaptive Targeted Dynamic Chunking for Tokenization-Free Hierarchical Model
by: Dang, Thang, et al.
Published: (2026)
by: Dang, Thang, et al.
Published: (2026)
CharBench: Evaluating the Role of Tokenization in Character-Level Tasks
by: Uzan, Omri, et al.
Published: (2025)
by: Uzan, Omri, et al.
Published: (2025)
The Strawberry Problem: Emergence of Character-level Understanding in Tokenized Language Models
by: Cosma, Adrian, et al.
Published: (2025)
by: Cosma, Adrian, et al.
Published: (2025)
Token Alignment via Character Matching for Subword Completion
by: Athiwaratkun, Ben, et al.
Published: (2024)
by: Athiwaratkun, Ben, et al.
Published: (2024)
Enhancing Character-Level Understanding in LLMs through Token Internal Structure Learning
by: Xu, Zhu, et al.
Published: (2024)
by: Xu, Zhu, et al.
Published: (2024)
Word Recovery in Large Language Models Enables Character-Level Tokenization Robustness
by: Yang, Zhipeng, et al.
Published: (2026)
by: Yang, Zhipeng, et al.
Published: (2026)
Empowering Character-level Text Infilling by Eliminating Sub-Tokens
by: Ren, Houxing, et al.
Published: (2024)
by: Ren, Houxing, et al.
Published: (2024)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
by: Shrestha, Adarsha, et al.
Published: (2025)
by: Shrestha, Adarsha, et al.
Published: (2025)
Broken-Token: Filtering Obfuscated Prompts by Counting Characters-Per-Token
by: Zychlinski, Shaked, et al.
Published: (2025)
by: Zychlinski, Shaked, et al.
Published: (2025)
Pretraining Language Models with Subword Regularization: An Empirical Study of BPE Dropout in Low-Resource NLP
by: Visser, Ruan, et al.
Published: (2026)
by: Visser, Ruan, et al.
Published: (2026)
KL3M Tokenizers: A Family of Domain-Specific and Character-Level Tokenizers for Legal, Financial, and Preprocessing Applications
by: Bommarito, Michael J, et al.
Published: (2025)
by: Bommarito, Michael J, et al.
Published: (2025)
Similar Items
-
Unifying Linear-Time Attention via Latent Probabilistic Modelling
by: Dolga, Rares, et al.
Published: (2024) -
Incremental Sequence Classification with Temporal Consistency
by: Maystre, Lucas, et al.
Published: (2025) -
When Embedding Models Meet: Procrustes Bounds and Applications
by: Maystre, Lucas, et al.
Published: (2025) -
BlockBPE: Parallel BPE Tokenization
by: You, Amos
Published: (2025) -
AdaptBPE: From General Purpose to Specialized Tokenizers
by: Liyanage, Vijini, et al.
Published: (2026)