Byte BPE Tokenization as an Inverse string Homomorphism
Fuente:
arXiv
Saved in:
| Main Authors: | Geng, Saibo, Gambhir, Sankalp, Wendler, Chris, West, Robert |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sketch-Guided Constrained Decoding for Boosting Blackbox Large Language Models without Logit Access
by: Geng, Saibo, et al.
Published: (2024)
by: Geng, Saibo, et al.
Published: (2024)
zip2zip: Inference-Time Adaptive Tokenization via Online Compression
by: Geng, Saibo, et al.
Published: (2025)
by: Geng, Saibo, et al.
Published: (2025)
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
by: Kadamba, Venu Gopal, et al.
Published: (2026)
by: Kadamba, Venu Gopal, et al.
Published: (2026)
BlockBPE: Parallel BPE Tokenization
by: You, Amos
Published: (2025)
by: You, Amos
Published: (2025)
Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning
by: Geng, Saibo, et al.
Published: (2023)
by: Geng, Saibo, et al.
Published: (2023)
Scaffold-BPE: Enhancing Byte Pair Encoding for Large Language Models with Simple and Effective Scaffold Token Removal
by: Lian, Haoran, et al.
Published: (2024)
by: Lian, Haoran, et al.
Published: (2024)
Batching BPE Tokenization Merges
by: Morgan, Alexander P.
Published: (2024)
by: Morgan, Alexander P.
Published: (2024)
TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards
by: Nica, Andreea, et al.
Published: (2025)
by: Nica, Andreea, et al.
Published: (2025)
Do Llamas Work in English? On the Latent Language of Multilingual Transformers
by: Wendler, Chris, et al.
Published: (2024)
by: Wendler, Chris, et al.
Published: (2024)
Constructing a BPE Tokenization DFA
by: Berglund, Martin, et al.
Published: (2024)
by: Berglund, Martin, et al.
Published: (2024)
LiteToken: Removing Intermediate Merge Residues From BPE Tokenizers
by: Sun, Yike, et al.
Published: (2026)
by: Sun, Yike, et al.
Published: (2026)
From Characters to Tokens: Dynamic Grouping with Hierarchical BPE
by: Dolga, Rares, et al.
Published: (2025)
by: Dolga, Rares, et al.
Published: (2025)
AdaptBPE: From General Purpose to Specialized Tokenizers
by: Liyanage, Vijini, et al.
Published: (2026)
by: Liyanage, Vijini, et al.
Published: (2026)
Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in Transformers
by: Dumas, Clément, et al.
Published: (2024)
by: Dumas, Clément, et al.
Published: (2024)
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
by: Chizhov, Pavel, et al.
Published: (2024)
by: Chizhov, Pavel, et al.
Published: (2024)
Adaptive BPE Tokenization for Enhanced Vocabulary Adaptation in Finetuning Pretrained Language Models
by: Balde, Gunjan, et al.
Published: (2024)
by: Balde, Gunjan, et al.
Published: (2024)
Evaluating Subword Tokenization Techniques for Bengali: A Benchmark Study with BengaliBPE
by: Patwary, Firoj Ahmmed, et al.
Published: (2025)
by: Patwary, Firoj Ahmmed, et al.
Published: (2025)
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
by: Hu, Yifan, et al.
Published: (2025)
by: Hu, Yifan, et al.
Published: (2025)
Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
by: Vemula, Saketh Reddy, et al.
Published: (2025)
by: Vemula, Saketh Reddy, et al.
Published: (2025)
JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models
by: Geng, Saibo, et al.
Published: (2025)
by: Geng, Saibo, et al.
Published: (2025)
Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?
by: Hayase, Jonathan, et al.
Published: (2024)
by: Hayase, Jonathan, et al.
Published: (2024)
Controllable Context Sensitivity and the Knob Behind It
by: Minder, Julian, et al.
Published: (2024)
by: Minder, Julian, et al.
Published: (2024)
Improbable Bigrams Expose Vulnerabilities of Incomplete Tokens in Byte-Level Tokenizers
by: Jang, Eugene, et al.
Published: (2024)
by: Jang, Eugene, et al.
Published: (2024)
Bit-level BPE: Below the byte boundary
by: Moon, Sangwhan, et al.
Published: (2025)
by: Moon, Sangwhan, et al.
Published: (2025)
Back to Bytes: Revisiting Tokenization Through UTF-8
by: Moryossef, Amit, et al.
Published: (2025)
by: Moryossef, Amit, et al.
Published: (2025)
Distilling Token-Trained Models into Byte-Level Models
by: Bao, Zishuo, et al.
Published: (2026)
by: Bao, Zishuo, et al.
Published: (2026)
Morphological Typology in BPE Subword Productivity and Language Modeling
by: Parra, Iñigo
Published: (2024)
by: Parra, Iñigo
Published: (2024)
An Analysis of BPE Vocabulary Trimming in Neural Machine Translation
by: Cognetta, Marco, et al.
Published: (2024)
by: Cognetta, Marco, et al.
Published: (2024)
MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies
by: Asgari, Ehsaneddin, et al.
Published: (2025)
by: Asgari, Ehsaneddin, et al.
Published: (2025)
ByteFlow: Language Modeling through Adaptive Byte Compression without a Tokenizer
by: Deng, Chunyuan, et al.
Published: (2026)
by: Deng, Chunyuan, et al.
Published: (2026)
Interpolation and Quantifiers in Ortholattices
by: Guilloud, Simon, et al.
Published: (2025)
by: Guilloud, Simon, et al.
Published: (2025)
LISA -- A Modern Proof System
by: Guilloud, Simon, et al.
Published: (2025)
by: Guilloud, Simon, et al.
Published: (2025)
Byte Latent Transformer: Patches Scale Better Than Tokens
by: Pagnoni, Artidoro, et al.
Published: (2024)
by: Pagnoni, Artidoro, et al.
Published: (2024)
SuperBPE: Space Travel for Language Models
by: Liu, Alisa, et al.
Published: (2025)
by: Liu, Alisa, et al.
Published: (2025)
BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization
by: Land, Sander, et al.
Published: (2025)
by: Land, Sander, et al.
Published: (2025)
Localized Cultural Knowledge is Conserved and Controllable in Large Language Models
by: Veselovsky, Veniamin, et al.
Published: (2025)
by: Veselovsky, Veniamin, et al.
Published: (2025)
Persona-Conditioned Risk Behavior in Large Language Models: A Simulated Gambling Study with GPT-4.1
by: Dubedy, Sankalp
Published: (2026)
by: Dubedy, Sankalp
Published: (2026)
Cross-Tokenizer LLM Distillation through a Byte-Level Interface
by: Singh, Avyav Kumar, et al.
Published: (2026)
by: Singh, Avyav Kumar, et al.
Published: (2026)
Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages
by: Brinkmann, Jannik, et al.
Published: (2025)
by: Brinkmann, Jannik, et al.
Published: (2025)
Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models
by: Sawada, Tomohiro, et al.
Published: (2025)
by: Sawada, Tomohiro, et al.
Published: (2025)
Similar Items
-
Sketch-Guided Constrained Decoding for Boosting Blackbox Large Language Models without Logit Access
by: Geng, Saibo, et al.
Published: (2024) -
zip2zip: Inference-Time Adaptive Tokenization via Online Compression
by: Geng, Saibo, et al.
Published: (2025) -
GPUTOK: GPU Accelerated Byte Level BPE Tokenization
by: Kadamba, Venu Gopal, et al.
Published: (2026) -
BlockBPE: Parallel BPE Tokenization
by: You, Amos
Published: (2025) -
Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning
by: Geng, Saibo, et al.
Published: (2023)