Investigating Continual Pretraining in Large Language Models: Insights and Implications
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yıldız, Çağatay, Ravichandran, Nishaanth Kanna, Sharma, Nitin, Bethge, Matthias, Ermis, Beyza |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
von: Öncel, Fırat, et al.
Veröffentlicht: (2024)
von: Öncel, Fırat, et al.
Veröffentlicht: (2024)
From Raw Corpora to Domain Benchmarks: Automated Evaluation of LLM Domain Expertise
von: Sharma, Nitin, et al.
Veröffentlicht: (2025)
von: Sharma, Nitin, et al.
Veröffentlicht: (2025)
From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models
von: Pozzobon, Luiza, et al.
Veröffentlicht: (2024)
von: Pozzobon, Luiza, et al.
Veröffentlicht: (2024)
Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs
von: Özer, Atahan, et al.
Veröffentlicht: (2025)
von: Özer, Atahan, et al.
Veröffentlicht: (2025)
On The Fairness Impacts of Hardware Selection in Machine Learning
von: Nelaturu, Sree Harsha, et al.
Veröffentlicht: (2023)
von: Nelaturu, Sree Harsha, et al.
Veröffentlicht: (2023)
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
Playing repeated games with Large Language Models
von: Akata, Elif, et al.
Veröffentlicht: (2023)
von: Akata, Elif, et al.
Veröffentlicht: (2023)
A Practitioner's Guide to Continual Multimodal Pretraining
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
von: Roth, Karsten, et al.
Veröffentlicht: (2024)
A Reproduction Study: The Kernel PCA Interpretation of Self-Attention Fails Under Scrutiny
von: Sarıtaş, Karahan, et al.
Veröffentlicht: (2025)
von: Sarıtaş, Karahan, et al.
Veröffentlicht: (2025)
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2024)
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2024)
Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning
von: Aakanksha, et al.
Veröffentlicht: (2024)
von: Aakanksha, et al.
Veröffentlicht: (2024)
Mapping Post-Training Forgetting in Language Models at Scale
von: Harmon, Jackson, et al.
Veröffentlicht: (2025)
von: Harmon, Jackson, et al.
Veröffentlicht: (2025)
Multilingual Arbitrage: Optimizing Data Pools to Accelerate Multilingual Progress
von: Odumakinde, Ayomide, et al.
Veröffentlicht: (2024)
von: Odumakinde, Ayomide, et al.
Veröffentlicht: (2024)
Emergent Abilities of Large Language Models under Continued Pretraining for Language Adaptation
von: Elhady, Ahmed, et al.
Veröffentlicht: (2025)
von: Elhady, Ahmed, et al.
Veröffentlicht: (2025)
The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models
von: Flamant, Cedric, et al.
Veröffentlicht: (2026)
von: Flamant, Cedric, et al.
Veröffentlicht: (2026)
Exploring the Benefits of Domain-Pretraining of Generative Large Language Models for Chemistry
von: Acharya, Anurag, et al.
Veröffentlicht: (2024)
von: Acharya, Anurag, et al.
Veröffentlicht: (2024)
Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
von: Stacey, Joe, et al.
Veröffentlicht: (2025)
von: Stacey, Joe, et al.
Veröffentlicht: (2025)
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
von: Skorobogat, Ronald, et al.
Veröffentlicht: (2026)
von: Skorobogat, Ronald, et al.
Veröffentlicht: (2026)
Object-level Self-Distillation for Vision Pretraining
von: Hızlı, Çağlar, et al.
Veröffentlicht: (2025)
von: Hızlı, Çağlar, et al.
Veröffentlicht: (2025)
Will Large Language Models Transform Clinical Prediction?
von: Yildiz, Yusuf, et al.
Veröffentlicht: (2025)
von: Yildiz, Yusuf, et al.
Veröffentlicht: (2025)
Infinite dSprites for Disentangled Continual Learning: Separating Memory Edits from Generalization
von: Dziadzio, Sebastian, et al.
Veröffentlicht: (2023)
von: Dziadzio, Sebastian, et al.
Veröffentlicht: (2023)
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
von: Hochlehnert, Andreas, et al.
Veröffentlicht: (2025)
von: Hochlehnert, Andreas, et al.
Veröffentlicht: (2025)
CiteME: Can Language Models Accurately Cite Scientific Claims?
von: Press, Ori, et al.
Veröffentlicht: (2024)
von: Press, Ori, et al.
Veröffentlicht: (2024)
Reuse, Don't Retrain: A Recipe for Continued Pretraining of Language Models
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
von: Parmar, Jupinder, et al.
Veröffentlicht: (2024)
MuCPT: Music-related Natural Language Model Continued Pretraining
von: Tian, Kai, et al.
Veröffentlicht: (2025)
von: Tian, Kai, et al.
Veröffentlicht: (2025)
The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining
von: Shao, Jiandong, et al.
Veröffentlicht: (2026)
von: Shao, Jiandong, et al.
Veröffentlicht: (2026)
Scalable Influence and Fact Tracing for Large Language Model Pretraining
von: Chang, Tyler A., et al.
Veröffentlicht: (2024)
von: Chang, Tyler A., et al.
Veröffentlicht: (2024)
Pretraining Exposure Explains Popularity Judgments in Large Language Models
von: Mozafari, Jamshid, et al.
Veröffentlicht: (2026)
von: Mozafari, Jamshid, et al.
Veröffentlicht: (2026)
Integrating Text and Time-Series into (Large) Language Models to Predict Medical Outcomes
von: Larbi, Iyadh Ben Cheikh, et al.
Veröffentlicht: (2025)
von: Larbi, Iyadh Ben Cheikh, et al.
Veröffentlicht: (2025)
PonderLM: Pretraining Language Models to Ponder in Continuous Space
von: Zeng, Boyi, et al.
Veröffentlicht: (2025)
von: Zeng, Boyi, et al.
Veröffentlicht: (2025)
Reflecting on the State of Rehearsal-free Continual Learning with Pretrained Models
von: Thede, Lukas, et al.
Veröffentlicht: (2024)
von: Thede, Lukas, et al.
Veröffentlicht: (2024)
Pretraining Large Language Models with NVFP4
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
MLP Memory: A Retriever-Pretrained Memory for Large Language Models
von: Wei, Rubin, et al.
Veröffentlicht: (2025)
von: Wei, Rubin, et al.
Veröffentlicht: (2025)
PretrainRL: Alleviating Factuality Hallucination of Large Language Models at the Beginning
von: Liu, Langming, et al.
Veröffentlicht: (2026)
von: Liu, Langming, et al.
Veröffentlicht: (2026)
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
von: Aakanksha, et al.
Veröffentlicht: (2024)
von: Aakanksha, et al.
Veröffentlicht: (2024)
From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
von: Rabbani, Parisa, et al.
Veröffentlicht: (2025)
von: Rabbani, Parisa, et al.
Veröffentlicht: (2025)
A Causal Language Modeling Detour Improves Encoder Continued Pretraining
von: Touchent, Rian, et al.
Veröffentlicht: (2026)
von: Touchent, Rian, et al.
Veröffentlicht: (2026)
AF Adapter: Continual Pretraining for Building Chinese Biomedical Language Model
von: Yan, Yongyu, et al.
Veröffentlicht: (2022)
von: Yan, Yongyu, et al.
Veröffentlicht: (2022)
DRAssist: Dispute Resolution Assistance using Large Language Models
von: Pawar, Sachin, et al.
Veröffentlicht: (2025)
von: Pawar, Sachin, et al.
Veröffentlicht: (2025)
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications
von: Tong, Ziyi, et al.
Veröffentlicht: (2026)
von: Tong, Ziyi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?
von: Öncel, Fırat, et al.
Veröffentlicht: (2024) -
From Raw Corpora to Domain Benchmarks: Automated Evaluation of LLM Domain Expertise
von: Sharma, Nitin, et al.
Veröffentlicht: (2025) -
From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models
von: Pozzobon, Luiza, et al.
Veröffentlicht: (2024) -
Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs
von: Özer, Atahan, et al.
Veröffentlicht: (2025) -
On The Fairness Impacts of Hardware Selection in Machine Learning
von: Nelaturu, Sree Harsha, et al.
Veröffentlicht: (2023)