Getting the most out of your tokenizer for pre-training and domain adaptation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dagan, Gautier, Synnaeve, Gabriel, Rozière, Baptiste |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Better & Faster Large Language Models via Multi-token Prediction
von: Gloeckle, Fabian, et al.
Veröffentlicht: (2024)
von: Gloeckle, Fabian, et al.
Veröffentlicht: (2024)
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
von: Chambon, Pierre, et al.
Veröffentlicht: (2025)
von: Chambon, Pierre, et al.
Veröffentlicht: (2025)
TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark
von: Jain, Kush, et al.
Veröffentlicht: (2024)
von: Jain, Kush, et al.
Veröffentlicht: (2024)
Meta Large Language Model Compiler: Foundation Models of Compiler Optimization
von: Cummins, Chris, et al.
Veröffentlicht: (2024)
von: Cummins, Chris, et al.
Veröffentlicht: (2024)
CAST: Cross-modal Alignment Similarity Test for Vision Language Models
von: Dagan, Gautier, et al.
Veröffentlicht: (2024)
von: Dagan, Gautier, et al.
Veröffentlicht: (2024)
Plancraft: an evaluation dataset for planning with LLM agents
von: Dagan, Gautier, et al.
Veröffentlicht: (2024)
von: Dagan, Gautier, et al.
Veröffentlicht: (2024)
$How^{2}$: How to learn from procedural How-to questions
von: Dagan, Gautier, et al.
Veröffentlicht: (2025)
von: Dagan, Gautier, et al.
Veröffentlicht: (2025)
Let your LLM generate a few tokens and you will reduce the need for retrieval
von: Déjean, Hervé
Veröffentlicht: (2024)
von: Déjean, Hervé
Veröffentlicht: (2024)
In-domain SSL pre-training and streaming ASR
von: Duret, Jarod, et al.
Veröffentlicht: (2025)
von: Duret, Jarod, et al.
Veröffentlicht: (2025)
Does your data spark joy? Performance gains from domain upsampling at the end of training
von: Blakeney, Cody, et al.
Veröffentlicht: (2024)
von: Blakeney, Cody, et al.
Veröffentlicht: (2024)
Do we really have to filter out random noise in pre-training data for language models?
von: Ru, Jinghan, et al.
Veröffentlicht: (2025)
von: Ru, Jinghan, et al.
Veröffentlicht: (2025)
Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
von: Hassid, Michael, et al.
Veröffentlicht: (2025)
von: Hassid, Michael, et al.
Veröffentlicht: (2025)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
von: Kumar, Anshul
Veröffentlicht: (2026)
von: Kumar, Anshul
Veröffentlicht: (2026)
Towards Nepali-language LLMs: Efficient GPT training with a Nepali BPE tokenizer
von: Shrestha, Adarsha, et al.
Veröffentlicht: (2025)
von: Shrestha, Adarsha, et al.
Veröffentlicht: (2025)
Adding LLMs to the psycholinguistic norming toolbox: A practical guide to getting the most out of human ratings
von: Conde, Javier, et al.
Veröffentlicht: (2025)
von: Conde, Javier, et al.
Veröffentlicht: (2025)
What Makes Large Language Models Reason in (Multi-Turn) Code Generation?
von: Zheng, Kunhao, et al.
Veröffentlicht: (2024)
von: Zheng, Kunhao, et al.
Veröffentlicht: (2024)
The KoLMogorov Test: Compression by Code Generation
von: Yoran, Ori, et al.
Veröffentlicht: (2025)
von: Yoran, Ori, et al.
Veröffentlicht: (2025)
Pre-training data selection for biomedical domain adaptation using journal impact metrics
von: Laï-king, Mathieu, et al.
Veröffentlicht: (2024)
von: Laï-king, Mathieu, et al.
Veröffentlicht: (2024)
Why do LLMs attend to the first token?
von: Barbero, Federico, et al.
Veröffentlicht: (2025)
von: Barbero, Federico, et al.
Veröffentlicht: (2025)
Where is the signal in tokenization space?
von: Geh, Renato Lui, et al.
Veröffentlicht: (2024)
von: Geh, Renato Lui, et al.
Veröffentlicht: (2024)
Interpreting token compositionality in LLMs: A robustness analysis
von: Aljaafari, Nura, et al.
Veröffentlicht: (2024)
von: Aljaafari, Nura, et al.
Veröffentlicht: (2024)
Comparative analysis of subword tokenization approaches for Indian languages
von: Das, Sudhansu Bala, et al.
Veröffentlicht: (2025)
von: Das, Sudhansu Bala, et al.
Veröffentlicht: (2025)
Contextual morphologically-guided tokenization for Latin encoder models
von: Hudspeth, Marisa, et al.
Veröffentlicht: (2025)
von: Hudspeth, Marisa, et al.
Veröffentlicht: (2025)
Beyond Pairwise: Global Zero-shot Temporal Graph Generation
von: Eirew, Alon, et al.
Veröffentlicht: (2025)
von: Eirew, Alon, et al.
Veröffentlicht: (2025)
AnomaLLMy -- Detecting anomalous tokens in black-box LLMs through low-confidence single-token predictions
von: Witold, Waligóra
Veröffentlicht: (2024)
von: Witold, Waligóra
Veröffentlicht: (2024)
Detecting harassment and defamation in cyberbullying with emotion-adaptive training
von: Yi, Peiling, et al.
Veröffentlicht: (2025)
von: Yi, Peiling, et al.
Veröffentlicht: (2025)
The pitfalls of next-token prediction
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage Retrieval
von: Ma, Guangyuan, et al.
Veröffentlicht: (2024)
von: Ma, Guangyuan, et al.
Veröffentlicht: (2024)
COVE: COntext and VEracity prediction for out-of-context images
von: Tonglet, Jonathan, et al.
Veröffentlicht: (2025)
von: Tonglet, Jonathan, et al.
Veröffentlicht: (2025)
On multi-token prediction for efficient LLM inference
von: Mehra, Somesh, et al.
Veröffentlicht: (2025)
von: Mehra, Somesh, et al.
Veröffentlicht: (2025)
Evaluating the performance of state-of-the-art esg domain-specific pre-trained large language models in text classification against existing models and traditional machine learning techniques
von: Chung, Tin Yuet, et al.
Veröffentlicht: (2024)
von: Chung, Tin Yuet, et al.
Veröffentlicht: (2024)
Continual Pre-training of MoEs: How robust is your router?
von: Thérien, Benjamin, et al.
Veröffentlicht: (2025)
von: Thérien, Benjamin, et al.
Veröffentlicht: (2025)
Collaborative decoding of critical tokens for boosting factuality of large language models
von: Jin, Lifeng, et al.
Veröffentlicht: (2024)
von: Jin, Lifeng, et al.
Veröffentlicht: (2024)
LBPE: Long-token-first Tokenization to Improve Large Language Models
von: Lian, Haoran, et al.
Veröffentlicht: (2024)
von: Lian, Haoran, et al.
Veröffentlicht: (2024)
Finetuning LLMs for EvaCun 2025 token prediction shared task
von: Jon, Josef, et al.
Veröffentlicht: (2025)
von: Jon, Josef, et al.
Veröffentlicht: (2025)
RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
von: Gehring, Jonas, et al.
Veröffentlicht: (2024)
von: Gehring, Jonas, et al.
Veröffentlicht: (2024)
Fine-tuning large language models for domain adaptation: Exploration of training strategies, scaling, model merging and synergistic capabilities
von: Lu, Wei, et al.
Veröffentlicht: (2024)
von: Lu, Wei, et al.
Veröffentlicht: (2024)
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
von: Singh, Aaditya K., et al.
Veröffentlicht: (2024)
von: Singh, Aaditya K., et al.
Veröffentlicht: (2024)
Do language models plan ahead for future tokens?
von: Wu, Wilson, et al.
Veröffentlicht: (2024)
von: Wu, Wilson, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Better & Faster Large Language Models via Multi-token Prediction
von: Gloeckle, Fabian, et al.
Veröffentlicht: (2024) -
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
von: Chambon, Pierre, et al.
Veröffentlicht: (2025) -
TestGenEval: A Real World Unit Test Generation and Test Completion Benchmark
von: Jain, Kush, et al.
Veröffentlicht: (2024) -
Meta Large Language Model Compiler: Foundation Models of Compiler Optimization
von: Cummins, Chris, et al.
Veröffentlicht: (2024) -
CAST: Cross-modal Alignment Similarity Test for Vision Language Models
von: Dagan, Gautier, et al.
Veröffentlicht: (2024)