Tokenizer Choice For LLM Training: Negligible or Crucial?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ali, Mehdi, Fromm, Michael, Thellmann, Klaudia, Rutmann, Richard, Lübbering, Max, Leveling, Johannes, Klug, Katrin, Ebert, Jan, Doll, Niclas, Buschhoff, Jasper Schulze, Jain, Charvi, Weber, Alexander Arno, Jurkschat, Lena, Abdelwahab, Hammam, John, Chelsea, Suarez, Pedro Ortiz, Ostendorff, Malte, Weinbach, Samuel, Sifa, Rafet, Kesselheim, Stefan, Flores-Herr, Nicolas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?
von: Doll, Niclas, et al.
Veröffentlicht: (2026)
von: Doll, Niclas, et al.
Veröffentlicht: (2026)
Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
von: Ali, Mehdi, et al.
Veröffentlicht: (2024)
von: Ali, Mehdi, et al.
Veröffentlicht: (2024)
Investigating Multilingual Instruction-Tuning: Do Polyglot Models Demand for Multilingual Instructions?
von: Weber, Alexander Arno, et al.
Veröffentlicht: (2024)
von: Weber, Alexander Arno, et al.
Veröffentlicht: (2024)
Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research
von: Lübbering, Max, et al.
Veröffentlicht: (2026)
von: Lübbering, Max, et al.
Veröffentlicht: (2026)
Towards Multilingual LLM Evaluation for European Languages
von: Thellmann, Klaudia, et al.
Veröffentlicht: (2024)
von: Thellmann, Klaudia, et al.
Veröffentlicht: (2024)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
von: Ali, Mehdi, et al.
Veröffentlicht: (2025)
von: Ali, Mehdi, et al.
Veröffentlicht: (2025)
Quantum Computing from Hopfield Nets
von: Bauckhage, Christian, et al.
Veröffentlicht: (2025)
von: Bauckhage, Christian, et al.
Veröffentlicht: (2025)
Diagnosing Translated Benchmarks: An Automated Quality Assurance Study of the EU20 Benchmark Suite
von: Thellmann, Klaudia, et al.
Veröffentlicht: (2026)
von: Thellmann, Klaudia, et al.
Veröffentlicht: (2026)
Data Processing for the OpenGPT-X Model Family
von: Brandizzi, Nicolo', et al.
Veröffentlicht: (2024)
von: Brandizzi, Nicolo', et al.
Veröffentlicht: (2024)
SynCED-EnDe 2025: A Synthetic and Curated English - German Dataset for Critical Error Detection in Machine Translation
von: Chopra, Muskaan, et al.
Veröffentlicht: (2025)
von: Chopra, Muskaan, et al.
Veröffentlicht: (2025)
Towards Reliable Machine Translation: Scaling LLMs for Critical Error Detection and Safety
von: Chopra, Muskaan, et al.
Veröffentlicht: (2026)
von: Chopra, Muskaan, et al.
Veröffentlicht: (2026)
Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation
von: Thellmann, Klaudia-Doris, et al.
Veröffentlicht: (2026)
von: Thellmann, Klaudia-Doris, et al.
Veröffentlicht: (2026)
Evaluation von Verwaltungsmodernisierung
von: Buschhoff, Christian
Veröffentlicht: (2020)
von: Buschhoff, Christian
Veröffentlicht: (2020)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
von: Penke, Carolin, et al.
Veröffentlicht: (2025)
von: Penke, Carolin, et al.
Veröffentlicht: (2025)
Model-agnostic Body Part Relevance Assessment for Pedestrian Detection
von: Günder, Maurice, et al.
Veröffentlicht: (2023)
von: Günder, Maurice, et al.
Veröffentlicht: (2023)
[Vision Paper] PRObot: Enhancing Patient-Reported Outcome Measures for Diabetic Retinopathy using Chatbots and Generative AI
von: Pielka, Maren, et al.
Veröffentlicht: (2024)
von: Pielka, Maren, et al.
Veröffentlicht: (2024)
Pointer-Guided Pre-Training: Infusing Large Language Models with Paragraph-Level Contextual Awareness
von: Hillebrand, Lars, et al.
Veröffentlicht: (2024)
von: Hillebrand, Lars, et al.
Veröffentlicht: (2024)
Reasoning LLMs in the Medical Domain: A Literature Survey
von: Berger, Armin, et al.
Veröffentlicht: (2025)
von: Berger, Armin, et al.
Veröffentlicht: (2025)
Interpretable Topic Extraction and Word Embedding Learning using row-stochastic DEDICOM
von: Hillebrand, Lars, et al.
Veröffentlicht: (2025)
von: Hillebrand, Lars, et al.
Veröffentlicht: (2025)
How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation
von: Chopra, Muskaan, et al.
Veröffentlicht: (2025)
von: Chopra, Muskaan, et al.
Veröffentlicht: (2025)
Knowing When Not to Predict: Self Supervised Learning and Abstention for Safer DR Screening
von: Chopra, Muskaan, et al.
Veröffentlicht: (2026)
von: Chopra, Muskaan, et al.
Veröffentlicht: (2026)
History Rhymes: Macro-Contextual Retrieval for Robust Financial Forecasting
von: Khanna, Sarthak, et al.
Veröffentlicht: (2025)
von: Khanna, Sarthak, et al.
Veröffentlicht: (2025)
Generalizing Abstention for Noise-Robust Learning in Medical Image Segmentation
von: Moustafa, Wesam, et al.
Veröffentlicht: (2026)
von: Moustafa, Wesam, et al.
Veröffentlicht: (2026)
Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing
von: Berghaus, David, et al.
Veröffentlicht: (2025)
von: Berghaus, David, et al.
Veröffentlicht: (2025)
Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit
von: Filatov, Oleg, et al.
Veröffentlicht: (2024)
von: Filatov, Oleg, et al.
Veröffentlicht: (2024)
Memory and Bandwidth are All You Need for Fully Sharded Data Parallel
von: Wang, Jiangtao, et al.
Veröffentlicht: (2025)
von: Wang, Jiangtao, et al.
Veröffentlicht: (2025)
Optimal Scaling Needs Optimal Norm
von: Filatov, Oleg, et al.
Veröffentlicht: (2025)
von: Filatov, Oleg, et al.
Veröffentlicht: (2025)
A Survey on Current Trends and Recent Advances in Text Anonymization
von: Deußer, Tobias, et al.
Veröffentlicht: (2025)
von: Deußer, Tobias, et al.
Veröffentlicht: (2025)
Towards Unified Multimodal Financial Forecasting: Integrating Sentiment Embeddings and Market Indicators via Cross-Modal Attention
von: Khanna, Sarthak, et al.
Veröffentlicht: (2025)
von: Khanna, Sarthak, et al.
Veröffentlicht: (2025)
When the Whole Is Less Than the Sum of Its Parts: Structural Coupling in Education
von: Raf Vanderstraeten, et al.
Veröffentlicht: (2026)
von: Raf Vanderstraeten, et al.
Veröffentlicht: (2026)
Negligence, Inadvertence, and Moral Responsibility: An Assessment of King’s ‘The Problem with Negligence’
von: Alejandro Mosqueda
Veröffentlicht: (2019)
von: Alejandro Mosqueda
Veröffentlicht: (2019)
The diverse roles of sphingolipids in inflammatory bowel disease
von: Chelsea L. Doll, et al.
Veröffentlicht: (2024)
von: Chelsea L. Doll, et al.
Veröffentlicht: (2024)
Is continuous CoT better suited for multi-lingual reasoning?
von: Bashir, Ali Hamza, et al.
Veröffentlicht: (2026)
von: Bashir, Ali Hamza, et al.
Veröffentlicht: (2026)
From Retinal Pixels to Patients: Evolution of Deep Learning Research in Diabetic Retinopathy Screening
von: Chopra, Muskaan, et al.
Veröffentlicht: (2025)
von: Chopra, Muskaan, et al.
Veröffentlicht: (2025)
Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks
von: Datta, Joyeeta, et al.
Veröffentlicht: (2025)
von: Datta, Joyeeta, et al.
Veröffentlicht: (2025)
Ausstellungskommunikation
von: Kesselheim, Wolfgang
Veröffentlicht: (2022)
von: Kesselheim, Wolfgang
Veröffentlicht: (2022)
Dataset for: Does Organized Misconduct Shape Network Topology? Part III: Topology of Retracted Article Citation Networks
von: IRMAK, Rafet
Veröffentlicht: (2026)
von: IRMAK, Rafet
Veröffentlicht: (2026)
Integración de pruebas y manejo de la restricción del crecimiento fetal
von: Sifa Turan
Veröffentlicht: (2008)
von: Sifa Turan
Veröffentlicht: (2008)
Abundance of Bitectatodinium spongium in surface sediments
von: Zonneveld, Karin A F, et al.
Veröffentlicht: (2002)
von: Zonneveld, Karin A F, et al.
Veröffentlicht: (2002)
HiStruct+: Improving Extractive Text Summarization with Hierarchical Structure Information
von: Ruan, Qian, et al.
Veröffentlicht: (2022)
von: Ruan, Qian, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?
von: Doll, Niclas, et al.
Veröffentlicht: (2026) -
Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
von: Ali, Mehdi, et al.
Veröffentlicht: (2024) -
Investigating Multilingual Instruction-Tuning: Do Polyglot Models Demand for Multilingual Instructions?
von: Weber, Alexander Arno, et al.
Veröffentlicht: (2024) -
Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research
von: Lübbering, Max, et al.
Veröffentlicht: (2026) -
Towards Multilingual LLM Evaluation for European Languages
von: Thellmann, Klaudia, et al.
Veröffentlicht: (2024)