Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation
Fuente:
arXiv
Salvato in:
| Autori principali: | Gigant, Théo, Peng, Bowen, Quesnelle, Jeffrey |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Pre-Training with Token Superposition
di: Peng, Bowen, et al.
Pubblicazione: (2026)
di: Peng, Bowen, et al.
Pubblicazione: (2026)
Long Context Pre-Training with Lighthouse Attention
di: Peng, Bowen, et al.
Pubblicazione: (2026)
di: Peng, Bowen, et al.
Pubblicazione: (2026)
Distilling Token-Trained Models into Byte-Level Models
di: Bao, Zishuo, et al.
Pubblicazione: (2026)
di: Bao, Zishuo, et al.
Pubblicazione: (2026)
YaRN: Efficient Context Window Extension of Large Language Models
di: Peng, Bowen, et al.
Pubblicazione: (2023)
di: Peng, Bowen, et al.
Pubblicazione: (2023)
Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure
di: Gigant, Théo, et al.
Pubblicazione: (2025)
di: Gigant, Théo, et al.
Pubblicazione: (2025)
Tokenization Falling Short: On Subword Robustness in Large Language Models
di: Chai, Yekun, et al.
Pubblicazione: (2024)
di: Chai, Yekun, et al.
Pubblicazione: (2024)
ByteSpan: Information-Driven Subword Tokenisation
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
di: Goriely, Zébulon, et al.
Pubblicazione: (2025)
Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge
di: Batsuren, Khuyagbaatar, et al.
Pubblicazione: (2024)
di: Batsuren, Khuyagbaatar, et al.
Pubblicazione: (2024)
Token Alignment via Character Matching for Subword Completion
di: Athiwaratkun, Ben, et al.
Pubblicazione: (2024)
di: Athiwaratkun, Ben, et al.
Pubblicazione: (2024)
Subword Tokenization Strategies for Kurdish Word Embeddings
di: Salehi, Ali, et al.
Pubblicazione: (2025)
di: Salehi, Ali, et al.
Pubblicazione: (2025)
Hermes 3 Technical Report
di: Teknium, Ryan, et al.
Pubblicazione: (2024)
di: Teknium, Ryan, et al.
Pubblicazione: (2024)
Understanding Subword Compositionality of Large Language Models
di: Peng, Qiwei, et al.
Pubblicazione: (2025)
di: Peng, Qiwei, et al.
Pubblicazione: (2025)
Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free Metrics
di: Gigant, Théo, et al.
Pubblicazione: (2024)
di: Gigant, Théo, et al.
Pubblicazione: (2024)
Evaluating Morphological Plausibility of Subword Tokenization via Statistical Alignment with Morpho-Syntactic Features
di: Stephen, Abishek, et al.
Pubblicazione: (2026)
di: Stephen, Abishek, et al.
Pubblicazione: (2026)
ByteFlow: Language Modeling through Adaptive Byte Compression without a Tokenizer
di: Deng, Chunyuan, et al.
Pubblicazione: (2026)
di: Deng, Chunyuan, et al.
Pubblicazione: (2026)
Morphological Typology in BPE Subword Productivity and Language Modeling
di: Parra, Iñigo
Pubblicazione: (2024)
di: Parra, Iñigo
Pubblicazione: (2024)
Assessing the Importance of Frequency versus Compositionality for Subword-based Tokenization in NMT
di: Wolleb, Benoist, et al.
Pubblicazione: (2023)
di: Wolleb, Benoist, et al.
Pubblicazione: (2023)
FuzzCoder: Byte-level Fuzzing Test via Large Language Model
di: Yang, Liqun, et al.
Pubblicazione: (2024)
di: Yang, Liqun, et al.
Pubblicazione: (2024)
MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
di: Kallini, Julie, et al.
Pubblicazione: (2024)
di: Kallini, Julie, et al.
Pubblicazione: (2024)
Evaluating Subword Tokenization Techniques for Bengali: A Benchmark Study with BengaliBPE
di: Patwary, Firoj Ahmmed, et al.
Pubblicazione: (2025)
di: Patwary, Firoj Ahmmed, et al.
Pubblicazione: (2025)
Team Ryu's Submission to SIGMORPHON 2024 Shared Task on Subword Tokenization
di: Li, Zilong
Pubblicazione: (2024)
di: Li, Zilong
Pubblicazione: (2024)
Scaffold-BPE: Enhancing Byte Pair Encoding for Large Language Models with Simple and Effective Scaffold Token Removal
di: Lian, Haoran, et al.
Pubblicazione: (2024)
di: Lian, Haoran, et al.
Pubblicazione: (2024)
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models
di: Zheng, Lin, et al.
Pubblicazione: (2026)
di: Zheng, Lin, et al.
Pubblicazione: (2026)
The Learning Dynamics of Subword Segmentation for Morphologically Diverse Languages
di: Meyer, Francois, et al.
Pubblicazione: (2025)
di: Meyer, Francois, et al.
Pubblicazione: (2025)
On the Effect of (Near) Duplicate Subwords in Language Modelling
di: Schäfer, Anton, et al.
Pubblicazione: (2024)
di: Schäfer, Anton, et al.
Pubblicazione: (2024)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
di: Deiseroth, Björn, et al.
Pubblicazione: (2024)
di: Deiseroth, Björn, et al.
Pubblicazione: (2024)
SpaceByte: Towards Deleting Tokenization from Large Language Modeling
di: Slagle, Kevin
Pubblicazione: (2024)
di: Slagle, Kevin
Pubblicazione: (2024)
Byte BPE Tokenization as an Inverse string Homomorphism
di: Geng, Saibo, et al.
Pubblicazione: (2024)
di: Geng, Saibo, et al.
Pubblicazione: (2024)
Distributional Properties of Subword Regularization
di: Cognetta, Marco, et al.
Pubblicazione: (2024)
di: Cognetta, Marco, et al.
Pubblicazione: (2024)
Lexically Grounded Subword Segmentation
di: Libovický, Jindřich, et al.
Pubblicazione: (2024)
di: Libovický, Jindřich, et al.
Pubblicazione: (2024)
Improbable Bigrams Expose Vulnerabilities of Incomplete Tokens in Byte-Level Tokenizers
di: Jang, Eugene, et al.
Pubblicazione: (2024)
di: Jang, Eugene, et al.
Pubblicazione: (2024)
Linguistic Laws Meet Protein Sequences: A Comparative Analysis of Subword Tokenization Methods
di: Suyunu, Burak, et al.
Pubblicazione: (2024)
di: Suyunu, Burak, et al.
Pubblicazione: (2024)
Hybrid Tokenization Strategy for DNA Language Model using Byte Pair Encoding and K-MER Methods
di: Sapkota, Ganesh, et al.
Pubblicazione: (2025)
di: Sapkota, Ganesh, et al.
Pubblicazione: (2025)
Exact Byte-Level Probabilities from Tokenized Language Models for FIM-Tasks and Model Ensembles
di: Phan, Buu, et al.
Pubblicazione: (2024)
di: Phan, Buu, et al.
Pubblicazione: (2024)
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
di: Hu, Yifan, et al.
Pubblicazione: (2025)
di: Hu, Yifan, et al.
Pubblicazione: (2025)
Back to Bytes: Revisiting Tokenization Through UTF-8
di: Moryossef, Amit, et al.
Pubblicazione: (2025)
di: Moryossef, Amit, et al.
Pubblicazione: (2025)
MambaByte: Token-free Selective State Space Model
di: Wang, Junxiong, et al.
Pubblicazione: (2024)
di: Wang, Junxiong, et al.
Pubblicazione: (2024)
Leading Whitespaces of Language Models' Subword Vocabulary Pose a Confound for Calculating Word Probabilities
di: Oh, Byung-Doh, et al.
Pubblicazione: (2024)
di: Oh, Byung-Doh, et al.
Pubblicazione: (2024)
Can Pretrained Language Models Derive Correct Semantics from Corrupt Subwords under Noise?
di: Li, Xinzhe, et al.
Pubblicazione: (2023)
di: Li, Xinzhe, et al.
Pubblicazione: (2023)
Stolen Subwords: Importance of Vocabularies for Machine Translation Model Stealing
di: Zouhar, Vilém
Pubblicazione: (2024)
di: Zouhar, Vilém
Pubblicazione: (2024)
Documenti analoghi
-
Efficient Pre-Training with Token Superposition
di: Peng, Bowen, et al.
Pubblicazione: (2026) -
Long Context Pre-Training with Lighthouse Attention
di: Peng, Bowen, et al.
Pubblicazione: (2026) -
Distilling Token-Trained Models into Byte-Level Models
di: Bao, Zishuo, et al.
Pubblicazione: (2026) -
YaRN: Efficient Context Window Extension of Large Language Models
di: Peng, Bowen, et al.
Pubblicazione: (2023) -
Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure
di: Gigant, Théo, et al.
Pubblicazione: (2025)