Tokenizer Choice For LLM Training: Negligible or Crucial?
Fuente:
arXiv
Saved in:
| Main Authors: | Ali, Mehdi, Fromm, Michael, Thellmann, Klaudia, Rutmann, Richard, Lübbering, Max, Leveling, Johannes, Klug, Katrin, Ebert, Jan, Doll, Niclas, Buschhoff, Jasper Schulze, Jain, Charvi, Weber, Alexander Arno, Jurkschat, Lena, Abdelwahab, Hammam, John, Chelsea, Suarez, Pedro Ortiz, Ostendorff, Malte, Weinbach, Samuel, Sifa, Rafet, Kesselheim, Stefan, Flores-Herr, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?
by: Doll, Niclas, et al.
Published: (2026)
by: Doll, Niclas, et al.
Published: (2026)
Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
by: Ali, Mehdi, et al.
Published: (2024)
by: Ali, Mehdi, et al.
Published: (2024)
Investigating Multilingual Instruction-Tuning: Do Polyglot Models Demand for Multilingual Instructions?
by: Weber, Alexander Arno, et al.
Published: (2024)
by: Weber, Alexander Arno, et al.
Published: (2024)
Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research
by: Lübbering, Max, et al.
Published: (2026)
by: Lübbering, Max, et al.
Published: (2026)
Towards Multilingual LLM Evaluation for European Languages
by: Thellmann, Klaudia, et al.
Published: (2024)
by: Thellmann, Klaudia, et al.
Published: (2024)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
by: Ali, Mehdi, et al.
Published: (2025)
by: Ali, Mehdi, et al.
Published: (2025)
Quantum Computing from Hopfield Nets
by: Bauckhage, Christian, et al.
Published: (2025)
by: Bauckhage, Christian, et al.
Published: (2025)
Diagnosing Translated Benchmarks: An Automated Quality Assurance Study of the EU20 Benchmark Suite
by: Thellmann, Klaudia, et al.
Published: (2026)
by: Thellmann, Klaudia, et al.
Published: (2026)
Data Processing for the OpenGPT-X Model Family
by: Brandizzi, Nicolo', et al.
Published: (2024)
by: Brandizzi, Nicolo', et al.
Published: (2024)
SynCED-EnDe 2025: A Synthetic and Curated English - German Dataset for Critical Error Detection in Machine Translation
by: Chopra, Muskaan, et al.
Published: (2025)
by: Chopra, Muskaan, et al.
Published: (2025)
Towards Reliable Machine Translation: Scaling LLMs for Critical Error Detection and Safety
by: Chopra, Muskaan, et al.
Published: (2026)
by: Chopra, Muskaan, et al.
Published: (2026)
Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation
by: Thellmann, Klaudia-Doris, et al.
Published: (2026)
by: Thellmann, Klaudia-Doris, et al.
Published: (2026)
Evaluation von Verwaltungsmodernisierung
by: Buschhoff, Christian
Published: (2020)
by: Buschhoff, Christian
Published: (2020)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
by: Penke, Carolin, et al.
Published: (2025)
by: Penke, Carolin, et al.
Published: (2025)
Model-agnostic Body Part Relevance Assessment for Pedestrian Detection
by: Günder, Maurice, et al.
Published: (2023)
by: Günder, Maurice, et al.
Published: (2023)
[Vision Paper] PRObot: Enhancing Patient-Reported Outcome Measures for Diabetic Retinopathy using Chatbots and Generative AI
by: Pielka, Maren, et al.
Published: (2024)
by: Pielka, Maren, et al.
Published: (2024)
Pointer-Guided Pre-Training: Infusing Large Language Models with Paragraph-Level Contextual Awareness
by: Hillebrand, Lars, et al.
Published: (2024)
by: Hillebrand, Lars, et al.
Published: (2024)
Reasoning LLMs in the Medical Domain: A Literature Survey
by: Berger, Armin, et al.
Published: (2025)
by: Berger, Armin, et al.
Published: (2025)
Interpretable Topic Extraction and Word Embedding Learning using row-stochastic DEDICOM
by: Hillebrand, Lars, et al.
Published: (2025)
by: Hillebrand, Lars, et al.
Published: (2025)
How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation
by: Chopra, Muskaan, et al.
Published: (2025)
by: Chopra, Muskaan, et al.
Published: (2025)
Knowing When Not to Predict: Self Supervised Learning and Abstention for Safer DR Screening
by: Chopra, Muskaan, et al.
Published: (2026)
by: Chopra, Muskaan, et al.
Published: (2026)
History Rhymes: Macro-Contextual Retrieval for Robust Financial Forecasting
by: Khanna, Sarthak, et al.
Published: (2025)
by: Khanna, Sarthak, et al.
Published: (2025)
Generalizing Abstention for Noise-Robust Learning in Medical Image Segmentation
by: Moustafa, Wesam, et al.
Published: (2026)
by: Moustafa, Wesam, et al.
Published: (2026)
Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing
by: Berghaus, David, et al.
Published: (2025)
by: Berghaus, David, et al.
Published: (2025)
Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit
by: Filatov, Oleg, et al.
Published: (2024)
by: Filatov, Oleg, et al.
Published: (2024)
Memory and Bandwidth are All You Need for Fully Sharded Data Parallel
by: Wang, Jiangtao, et al.
Published: (2025)
by: Wang, Jiangtao, et al.
Published: (2025)
Optimal Scaling Needs Optimal Norm
by: Filatov, Oleg, et al.
Published: (2025)
by: Filatov, Oleg, et al.
Published: (2025)
A Survey on Current Trends and Recent Advances in Text Anonymization
by: Deußer, Tobias, et al.
Published: (2025)
by: Deußer, Tobias, et al.
Published: (2025)
Towards Unified Multimodal Financial Forecasting: Integrating Sentiment Embeddings and Market Indicators via Cross-Modal Attention
by: Khanna, Sarthak, et al.
Published: (2025)
by: Khanna, Sarthak, et al.
Published: (2025)
When the Whole Is Less Than the Sum of Its Parts: Structural Coupling in Education
by: Raf Vanderstraeten, et al.
Published: (2026)
by: Raf Vanderstraeten, et al.
Published: (2026)
Negligence, Inadvertence, and Moral Responsibility: An Assessment of King’s ‘The Problem with Negligence’
by: Alejandro Mosqueda
Published: (2019)
by: Alejandro Mosqueda
Published: (2019)
The diverse roles of sphingolipids in inflammatory bowel disease
by: Chelsea L. Doll, et al.
Published: (2024)
by: Chelsea L. Doll, et al.
Published: (2024)
Is continuous CoT better suited for multi-lingual reasoning?
by: Bashir, Ali Hamza, et al.
Published: (2026)
by: Bashir, Ali Hamza, et al.
Published: (2026)
From Retinal Pixels to Patients: Evolution of Deep Learning Research in Diabetic Retinopathy Screening
by: Chopra, Muskaan, et al.
Published: (2025)
by: Chopra, Muskaan, et al.
Published: (2025)
Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks
by: Datta, Joyeeta, et al.
Published: (2025)
by: Datta, Joyeeta, et al.
Published: (2025)
Ausstellungskommunikation
by: Kesselheim, Wolfgang
Published: (2022)
by: Kesselheim, Wolfgang
Published: (2022)
Dataset for: Does Organized Misconduct Shape Network Topology? Part III: Topology of Retracted Article Citation Networks
by: IRMAK, Rafet
Published: (2026)
by: IRMAK, Rafet
Published: (2026)
Integración de pruebas y manejo de la restricción del crecimiento fetal
by: Sifa Turan
Published: (2008)
by: Sifa Turan
Published: (2008)
Abundance of Bitectatodinium spongium in surface sediments
by: Zonneveld, Karin A F, et al.
Published: (2002)
by: Zonneveld, Karin A F, et al.
Published: (2002)
HiStruct+: Improving Extractive Text Summarization with Hierarchical Structure Information
by: Ruan, Qian, et al.
Published: (2022)
by: Ruan, Qian, et al.
Published: (2022)
Similar Items
-
Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?
by: Doll, Niclas, et al.
Published: (2026) -
Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
by: Ali, Mehdi, et al.
Published: (2024) -
Investigating Multilingual Instruction-Tuning: Do Polyglot Models Demand for Multilingual Instructions?
by: Weber, Alexander Arno, et al.
Published: (2024) -
Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research
by: Lübbering, Max, et al.
Published: (2026) -
Towards Multilingual LLM Evaluation for European Languages
by: Thellmann, Klaudia, et al.
Published: (2024)