Simple and Scalable Strategies to Continually Pre-train Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ibrahim, Adam, Thérien, Benjamin, Gupta, Kshitij, Richter, Mats L., Anthony, Quentin, Lesort, Timothée, Belilovsky, Eugene, Rish, Irina |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025)
Continual Pre-training of MoEs: How robust is your router?
von: Thérien, Benjamin, et al.
Veröffentlicht: (2025)
von: Thérien, Benjamin, et al.
Veröffentlicht: (2025)
MuLoCo: Muon is a practical inner optimizer for DiLoCo
von: Thérien, Benjamin, et al.
Veröffentlicht: (2025)
von: Thérien, Benjamin, et al.
Veröffentlicht: (2025)
Communication Efficient LLM Pre-training with SparseLoCo
von: Sarfi, Amir, et al.
Veröffentlicht: (2025)
von: Sarfi, Amir, et al.
Veröffentlicht: (2025)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
von: Legate, Gwen, et al.
Veröffentlicht: (2025)
von: Legate, Gwen, et al.
Veröffentlicht: (2025)
Continual Learning Under Language Shift
von: Gogoulou, Evangelia, et al.
Veröffentlicht: (2023)
von: Gogoulou, Evangelia, et al.
Veröffentlicht: (2023)
$μ$LO: Compute-Efficient Meta-Generalization of Learned Optimizers
von: Thérien, Benjamin, et al.
Veröffentlicht: (2024)
von: Thérien, Benjamin, et al.
Veröffentlicht: (2024)
PyLO: Towards Accessible Learned Optimizers in PyTorch
von: Janson, Paul, et al.
Veröffentlicht: (2025)
von: Janson, Paul, et al.
Veröffentlicht: (2025)
Covenant-72B: Pre-Training a 72B LLM with Trustless Peers Over-the-Internet
von: Lidin, Joel, et al.
Veröffentlicht: (2026)
von: Lidin, Joel, et al.
Veröffentlicht: (2026)
From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for Tibetan
von: Yang, Lei, et al.
Veröffentlicht: (2025)
von: Yang, Lei, et al.
Veröffentlicht: (2025)
Amplifying Pathological Detection in EEG Signaling Pathways through Cross-Dataset Transfer Learning
von: Darvishi-Bayazi, Mohammad-Javad, et al.
Veröffentlicht: (2023)
von: Darvishi-Bayazi, Mohammad-Javad, et al.
Veröffentlicht: (2023)
Examining Forgetting in Continual Pre-training of Aligned Large Language Models
von: Li, Chen-An, et al.
Veröffentlicht: (2024)
von: Li, Chen-An, et al.
Veröffentlicht: (2024)
Meta-learning Optimizers for Communication-Efficient Learning
von: Joseph, Charles-Étienne, et al.
Veröffentlicht: (2023)
von: Joseph, Charles-Étienne, et al.
Veröffentlicht: (2023)
Pre-trained Language Models Learn Remarkably Accurate Representations of Numbers
von: Kadlčík, Marek, et al.
Veröffentlicht: (2025)
von: Kadlčík, Marek, et al.
Veröffentlicht: (2025)
WHODUNIT: Evaluation benchmark for culprit detection in mystery stories
von: Gupta, Kshitij
Veröffentlicht: (2025)
von: Gupta, Kshitij
Veröffentlicht: (2025)
Efficient Continual Pre-training for Building Domain Specific Large Language Models
von: Xie, Yong, et al.
Veröffentlicht: (2023)
von: Xie, Yong, et al.
Veröffentlicht: (2023)
Towards Effective and Efficient Continual Pre-training of Large Language Models
von: Chen, Jie, et al.
Veröffentlicht: (2024)
von: Chen, Jie, et al.
Veröffentlicht: (2024)
Aurora-M: Open Source Continual Pre-training for Multilingual Language and Code
von: Nakamura, Taishi, et al.
Veröffentlicht: (2024)
von: Nakamura, Taishi, et al.
Veröffentlicht: (2024)
A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification
von: Calbucura, Nicolas, et al.
Veröffentlicht: (2025)
von: Calbucura, Nicolas, et al.
Veröffentlicht: (2025)
HLAT: High-quality Large Language Model Pre-trained on AWS Trainium
von: Fan, Haozheng, et al.
Veröffentlicht: (2024)
von: Fan, Haozheng, et al.
Veröffentlicht: (2024)
Construction of Domain-specified Japanese Large Language Model for Finance through Continual Pre-training
von: Hirano, Masanori, et al.
Veröffentlicht: (2024)
von: Hirano, Masanori, et al.
Veröffentlicht: (2024)
Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models
von: Ma, Shengjie, et al.
Veröffentlicht: (2025)
von: Ma, Shengjie, et al.
Veröffentlicht: (2025)
Efficient Continual Pre-training of LLMs for Low-resource Languages
von: Nag, Arijit, et al.
Veröffentlicht: (2024)
von: Nag, Arijit, et al.
Veröffentlicht: (2024)
Model Merging in Pre-training of Large Language Models
von: Li, Yunshui, et al.
Veröffentlicht: (2025)
von: Li, Yunshui, et al.
Veröffentlicht: (2025)
NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data
von: Bogdanov, Sergei, et al.
Veröffentlicht: (2024)
von: Bogdanov, Sergei, et al.
Veröffentlicht: (2024)
Scaling Agents via Continual Pre-training
von: Su, Liangcai, et al.
Veröffentlicht: (2025)
von: Su, Liangcai, et al.
Veröffentlicht: (2025)
Zyda: A 1.3T Dataset for Open Language Modeling
von: Tokpanov, Yury, et al.
Veröffentlicht: (2024)
von: Tokpanov, Yury, et al.
Veröffentlicht: (2024)
STEP: Staged Parameter-Efficient Pre-training for Large Language Models
von: Yano, Kazuki, et al.
Veröffentlicht: (2025)
von: Yano, Kazuki, et al.
Veröffentlicht: (2025)
Cross-layer Attention Sharing for Pre-trained Large Language Models
von: Mu, Yongyu, et al.
Veröffentlicht: (2024)
von: Mu, Yongyu, et al.
Veröffentlicht: (2024)
Pre-trained Large Language Models for Financial Sentiment Analysis
von: Luo, Wei, et al.
Veröffentlicht: (2024)
von: Luo, Wei, et al.
Veröffentlicht: (2024)
Spike No More: Stabilizing the Pre-training of Large Language Models
von: Takase, Sho, et al.
Veröffentlicht: (2023)
von: Takase, Sho, et al.
Veröffentlicht: (2023)
RedPajama: an Open Dataset for Training Large Language Models
von: Weber, Maurice, et al.
Veröffentlicht: (2024)
von: Weber, Maurice, et al.
Veröffentlicht: (2024)
Efficient Continual Pre-training by Mitigating the Stability Gap
von: Guo, Yiduo, et al.
Veröffentlicht: (2024)
von: Guo, Yiduo, et al.
Veröffentlicht: (2024)
Bag of Lies: Robustness in Continuous Pre-training BERT
von: Gevers, Ine, et al.
Veröffentlicht: (2024)
von: Gevers, Ine, et al.
Veröffentlicht: (2024)
Enhancing Question Answering on Charts Through Effective Pre-training Tasks
von: Gupta, Ashim, et al.
Veröffentlicht: (2024)
von: Gupta, Ashim, et al.
Veröffentlicht: (2024)
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
von: Liu, Hong, et al.
Veröffentlicht: (2023)
von: Liu, Hong, et al.
Veröffentlicht: (2023)
Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models
von: Tu, Zhijun, et al.
Veröffentlicht: (2025)
von: Tu, Zhijun, et al.
Veröffentlicht: (2025)
Language Portability Strategies for Open-domain Dialogue with Pre-trained Language Models from High to Low Resource Languages
von: Njifenjou, Ahmed, et al.
Veröffentlicht: (2024)
von: Njifenjou, Ahmed, et al.
Veröffentlicht: (2024)
Machine Unlearning of Pre-trained Large Language Models
von: Yao, Jin, et al.
Veröffentlicht: (2024)
von: Yao, Jin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
von: Singh, Vaibhav, et al.
Veröffentlicht: (2025) -
Continual Pre-training of MoEs: How robust is your router?
von: Thérien, Benjamin, et al.
Veröffentlicht: (2025) -
MuLoCo: Muon is a practical inner optimizer for DiLoCo
von: Thérien, Benjamin, et al.
Veröffentlicht: (2025) -
Communication Efficient LLM Pre-training with SparseLoCo
von: Sarfi, Amir, et al.
Veröffentlicht: (2025) -
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)