From Where Words Come: Efficient Regularization of Code Tokenizers Through Source Attribution
Fuente:
arXiv
Saved in:
| Main Authors: | Chizhov, Pavel, Bogomolov, Egor, Yamshchikov, Ivan P. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models
by: Purason, Taido, et al.
Published: (2025)
by: Purason, Taido, et al.
Published: (2025)
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
by: Chizhov, Pavel, et al.
Published: (2024)
by: Chizhov, Pavel, et al.
Published: (2024)
Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models
by: Sorokovikova, Aleksandra, et al.
Published: (2025)
by: Sorokovikova, Aleksandra, et al.
Published: (2025)
What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks
by: Chizhov, Pavel, et al.
Published: (2025)
by: Chizhov, Pavel, et al.
Published: (2025)
The Company You Keep: How LLMs Respond to Dark Triad Traits
by: Lu, Zeyi, et al.
Published: (2026)
by: Lu, Zeyi, et al.
Published: (2026)
Model in Distress: Sentiment Analysis on French Synthetic Social Media
by: Langlais, Pierre-Carl, et al.
Published: (2026)
by: Langlais, Pierre-Carl, et al.
Published: (2026)
Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family
by: Langlais, Pierre-Carl, et al.
Published: (2025)
by: Langlais, Pierre-Carl, et al.
Published: (2025)
Sui Generis: Large Language Models for Authorship Attribution and Verification in Latin
by: Schmidt, Gleb, et al.
Published: (2024)
by: Schmidt, Gleb, et al.
Published: (2024)
Knowledge Graph Representation for Political Information Sources
by: Osmonova, Tinatin, et al.
Published: (2024)
by: Osmonova, Tinatin, et al.
Published: (2024)
Vocabulary Transfer for Biomedical Texts: Add Tokens if You Can Not Add Data
by: Singh, Priyanka, et al.
Published: (2022)
by: Singh, Priyanka, et al.
Published: (2022)
Toxicity of the Commons: Curating Open-Source Pre-Training Data
by: Arnett, Catherine, et al.
Published: (2024)
by: Arnett, Catherine, et al.
Published: (2024)
Transfer of Structural Knowledge from Synthetic Languages
by: Budnikov, Mikhail, et al.
Published: (2025)
by: Budnikov, Mikhail, et al.
Published: (2025)
Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
by: Langlais, Pierre-Carl, et al.
Published: (2025)
by: Langlais, Pierre-Carl, et al.
Published: (2025)
ComicScene154: A Scene Dataset for Comic Analysis
by: Paval, Sandro, et al.
Published: (2025)
by: Paval, Sandro, et al.
Published: (2025)
Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
by: Liu, Kaiyuan, et al.
Published: (2025)
by: Liu, Kaiyuan, et al.
Published: (2025)
Understanding LLMs' Cross-Lingual Context Retrieval: How Good It Is And Where It Comes From
by: Gao, Changjiang, et al.
Published: (2025)
by: Gao, Changjiang, et al.
Published: (2025)
Do Data-based Curricula Work?
by: Surkov, Maxim K., et al.
Published: (2021)
by: Surkov, Maxim K., et al.
Published: (2021)
What is Wrong with Language Models that Can Not Tell a Story?
by: Yamshchikov, Ivan P., et al.
Published: (2022)
by: Yamshchikov, Ivan P., et al.
Published: (2022)
Step Rejection Fine-Tuning: A Practical Distillation Recipe
by: Slinko, Igor, et al.
Published: (2026)
by: Slinko, Igor, et al.
Published: (2026)
On the Semantic and Syntactic Information Encoded in Proto-Tokens for One-Step Text Reconstruction
by: Bondarenko, Ivan, et al.
Published: (2026)
by: Bondarenko, Ivan, et al.
Published: (2026)
From Tokens to Words: On the Inner Lexicon of LLMs
by: Kaplan, Guy, et al.
Published: (2024)
by: Kaplan, Guy, et al.
Published: (2024)
VISTA: Visualization of Token Attribution via Efficient Analysis
by: Ahmed, Syed, et al.
Published: (2026)
by: Ahmed, Syed, et al.
Published: (2026)
Back to Bytes: Revisiting Tokenization Through UTF-8
by: Moryossef, Amit, et al.
Published: (2025)
by: Moryossef, Amit, et al.
Published: (2025)
CleanComedy: Creating Friendly Humor through Generative Techniques
by: Vikhorev, Dmitry, et al.
Published: (2024)
by: Vikhorev, Dmitry, et al.
Published: (2024)
Better Call Claude: Can LLMs Detect Changes of Writing Style?
by: Römisch, Johannes, et al.
Published: (2025)
by: Römisch, Johannes, et al.
Published: (2025)
Team "better_call_claude": Style Change Detection using a Sequential Sentence Pair Classifier
by: Schmidt, Gleb, et al.
Published: (2025)
by: Schmidt, Gleb, et al.
Published: (2025)
Neural Machine Translation for Malayalam Paraphrase Generation
by: Varghese, Christeena, et al.
Published: (2024)
by: Varghese, Christeena, et al.
Published: (2024)
Subword Tokenization Strategies for Kurdish Word Embeddings
by: Salehi, Ali, et al.
Published: (2025)
by: Salehi, Ali, et al.
Published: (2025)
Echo-chambers and Idea Labs: Communication Styles on Twitter
by: Sorokovikova, Aleksandra, et al.
Published: (2024)
by: Sorokovikova, Aleksandra, et al.
Published: (2024)
On Problems of Implicit Context Compression for Software Engineering Agents
by: Gelvan, Kirill, et al.
Published: (2026)
by: Gelvan, Kirill, et al.
Published: (2026)
TokenShapley: Token Level Context Attribution with Shapley Value
by: Xiao, Yingtai, et al.
Published: (2025)
by: Xiao, Yingtai, et al.
Published: (2025)
Broken Words, Broken Performance: Effect of Tokenization on Performance of LLMs
by: Pawar, Sachin, et al.
Published: (2025)
by: Pawar, Sachin, et al.
Published: (2025)
Individuation in Neural Models with and without Visual Grounding
by: Tikhonov, Alexey, et al.
Published: (2024)
by: Tikhonov, Alexey, et al.
Published: (2024)
Improving Small Language Models for Code Generation with Reinforcement Learning from Verification Feedback
by: Skopin, Egor, et al.
Published: (2026)
by: Skopin, Egor, et al.
Published: (2026)
EfficientXLang: Towards Improving Token Efficiency Through Cross-Lingual Reasoning
by: Ahuja, Sanchit, et al.
Published: (2025)
by: Ahuja, Sanchit, et al.
Published: (2025)
From Token to Line: Enhancing Code Generation with a Long-Term Perspective
by: Lu, Tingwei, et al.
Published: (2025)
by: Lu, Tingwei, et al.
Published: (2025)
LLMs Simulate Big Five Personality Traits: Further Evidence
by: Sorokovikova, Aleksandra, et al.
Published: (2024)
by: Sorokovikova, Aleksandra, et al.
Published: (2024)
Parsing Through Boundaries in Chinese Word Segmentation
by: Chen, Yige, et al.
Published: (2025)
by: Chen, Yige, et al.
Published: (2025)
The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
by: Thakur, Aamod, et al.
Published: (2025)
by: Thakur, Aamod, et al.
Published: (2025)
CROP: Token-Efficient Reasoning in Large Language Models via Regularized Prompt Optimization
by: Shah, Deep, et al.
Published: (2026)
by: Shah, Deep, et al.
Published: (2026)
Similar Items
-
Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models
by: Purason, Taido, et al.
Published: (2025) -
BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training
by: Chizhov, Pavel, et al.
Published: (2024) -
Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models
by: Sorokovikova, Aleksandra, et al.
Published: (2025) -
What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks
by: Chizhov, Pavel, et al.
Published: (2025) -
The Company You Keep: How LLMs Respond to Dark Triad Traits
by: Lu, Zeyi, et al.
Published: (2026)