Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Abbes, Istabrak, Subbaraj, Gopeshh, Riemer, Matthew, Islah, Nizar, Therien, Benjamin, Tabaru, Tsuguchika, Kingetsu, Hiroaki, Chandar, Sarath, Rish, Irina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference
by: Riemer, Matthew, et al.
Published: (2024)
by: Riemer, Matthew, et al.
Published: (2024)
Small Encoders Can Rival Large Decoders in Detecting Groundedness
by: Abbes, Istabrak, et al.
Published: (2025)
by: Abbes, Istabrak, et al.
Published: (2025)
GFlowNet Pretraining with Inexpensive Rewards
by: Pandey, Mohit, et al.
Published: (2024)
by: Pandey, Mohit, et al.
Published: (2024)
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning
by: Arefin, Md Rifat, et al.
Published: (2024)
by: Arefin, Md Rifat, et al.
Published: (2024)
Direct Quantized Training of Language Models with Stochastic Rounding
by: Zhao, Kaiyan, et al.
Published: (2024)
by: Zhao, Kaiyan, et al.
Published: (2024)
Simple and Scalable Strategies to Continually Pre-train Large Language Models
by: Ibrahim, Adam, et al.
Published: (2024)
by: Ibrahim, Adam, et al.
Published: (2024)
MuLoCo: Muon is a practical inner optimizer for DiLoCo
by: Thérien, Benjamin, et al.
Published: (2025)
by: Thérien, Benjamin, et al.
Published: (2025)
GitChameleon: Unmasking the Version-Switching Capabilities of Code Generation Models
by: Islah, Nizar, et al.
Published: (2024)
by: Islah, Nizar, et al.
Published: (2024)
Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs
by: Thakkar, Megh, et al.
Published: (2024)
by: Thakkar, Megh, et al.
Published: (2024)
Exploring Quantization for Efficient Pre-Training of Transformer Language Models
by: Chitsaz, Kamran, et al.
Published: (2024)
by: Chitsaz, Kamran, et al.
Published: (2024)
A Deep Dive into the Trade-Offs of Parameter-Efficient Preference Alignment Techniques
by: Thakkar, Megh, et al.
Published: (2024)
by: Thakkar, Megh, et al.
Published: (2024)
Continual Pre-training of MoEs: How robust is your router?
by: Thérien, Benjamin, et al.
Published: (2025)
by: Thérien, Benjamin, et al.
Published: (2025)
Pretraining Generative Flow Networks with Inexpensive Rewards for Molecular Graph Generation
by: Pandey, Mohit, et al.
Published: (2025)
by: Pandey, Mohit, et al.
Published: (2025)
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
by: Legate, Gwen, et al.
Published: (2025)
by: Legate, Gwen, et al.
Published: (2025)
$μ$LO: Compute-Efficient Meta-Generalization of Learned Optimizers
by: Thérien, Benjamin, et al.
Published: (2024)
by: Thérien, Benjamin, et al.
Published: (2024)
Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training
by: Ostapenko, Oleksiy, et al.
Published: (2025)
by: Ostapenko, Oleksiy, et al.
Published: (2025)
Handling Delay in Real-Time Reinforcement Learning
by: Anokhin, Ivan, et al.
Published: (2025)
by: Anokhin, Ivan, et al.
Published: (2025)
Towards Practical Tool Usage for Continually Learning LLMs
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
Faithfulness Measurable Masked Language Models
by: Madsen, Andreas, et al.
Published: (2023)
by: Madsen, Andreas, et al.
Published: (2023)
Are self-explanations from Large Language Models faithful?
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing
by: Hashemzadeh, Maryam, et al.
Published: (2026)
by: Hashemzadeh, Maryam, et al.
Published: (2026)
Revisiting Softmax Masking: Stop Gradient for Enhancing Stability in Replay-based Continual Learning
by: Kim, Hoyong, et al.
Published: (2023)
by: Kim, Hoyong, et al.
Published: (2023)
Communication Efficient LLM Pre-training with SparseLoCo
by: Sarfi, Amir, et al.
Published: (2025)
by: Sarfi, Amir, et al.
Published: (2025)
CoPeP: Benchmarking Continual Pretraining for Protein Language Models
by: Patil, Darshan, et al.
Published: (2026)
by: Patil, Darshan, et al.
Published: (2026)
The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
by: Riemer, Matthew, et al.
Published: (2025)
by: Riemer, Matthew, et al.
Published: (2025)
Efficacy of Topical Nigella sativa Oil for Oral Wound Healing in Rats
by: Alper Tabaru, et al.
Published: (2025)
by: Alper Tabaru, et al.
Published: (2025)
GitChameleon 2.0: Evaluating AI Code Generation Against Python Library Version Incompatibilities
by: Misra, Diganta, et al.
Published: (2025)
by: Misra, Diganta, et al.
Published: (2025)
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
by: Nilaksh, et al.
Published: (2026)
by: Nilaksh, et al.
Published: (2026)
NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
by: Bhan, Milan, et al.
Published: (2025)
by: Bhan, Milan, et al.
Published: (2025)
Learning to combine top-down context and feed-forward representations under ambiguity with apical and basal dendrites
by: Islah, Nizar, et al.
Published: (2023)
by: Islah, Nizar, et al.
Published: (2023)
LLMs Can't Play Hangman: On the Necessity of a Private Working Memory for Language Agents
by: Baldelli, Davide, et al.
Published: (2026)
by: Baldelli, Davide, et al.
Published: (2026)
Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
by: Guiroy, Simon, et al.
Published: (2025)
by: Guiroy, Simon, et al.
Published: (2025)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
by: Prato, Gabriele, et al.
Published: (2025)
by: Prato, Gabriele, et al.
Published: (2025)
The Expressive Limits of Diagonal SSMs for State-Tracking
by: Shakerinava, Mehran, et al.
Published: (2026)
by: Shakerinava, Mehran, et al.
Published: (2026)
Intelligent Switching for Reset-Free RL
by: Patil, Darshan, et al.
Published: (2024)
by: Patil, Darshan, et al.
Published: (2024)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
Interpretability Needs a New Paradigm
by: Madsen, Andreas, et al.
Published: (2024)
by: Madsen, Andreas, et al.
Published: (2024)
Why Don't Prompt-Based Fairness Metrics Correlate?
by: Zayed, Abdelrahman, et al.
Published: (2024)
by: Zayed, Abdelrahman, et al.
Published: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
by: Zayed, Abdelrahman, et al.
Published: (2023)
by: Zayed, Abdelrahman, et al.
Published: (2023)
Similar Items
-
Enabling Realtime Reinforcement Learning at Scale with Staggered Asynchronous Inference
by: Riemer, Matthew, et al.
Published: (2024) -
Small Encoders Can Rival Large Decoders in Detecting Groundedness
by: Abbes, Istabrak, et al.
Published: (2025) -
GFlowNet Pretraining with Inexpensive Rewards
by: Pandey, Mohit, et al.
Published: (2024) -
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
by: Singh, Vaibhav, et al.
Published: (2025) -
Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning
by: Arefin, Md Rifat, et al.
Published: (2024)