Logits Replay + MoClip: Stabilized, Low-Cost Post-Training with Minimal Forgetting
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Suming, Li, Jing, Zhou, Zhicheng, Huang, Junjie, Qiu, Linyuan, Sun, Zhijie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data Trajectory Alignment for LLM Domain Adaptation: A Two-Phase Synthesis Framework for Telecommunications Mathematics
by: Zhou, Zhicheng, et al.
Published: (2025)
by: Zhou, Zhicheng, et al.
Published: (2025)
Agentic-KGR: Co-evolutionary Knowledge Graph Construction through Multi-Agent Reinforcement Learning
by: Li, Jing, et al.
Published: (2025)
by: Li, Jing, et al.
Published: (2025)
DeepJSONEval: Benchmarking Complex Nested JSON Data Mining for Large Language Models
by: Zhou, Zhicheng, et al.
Published: (2025)
by: Zhou, Zhicheng, et al.
Published: (2025)
HES-SQL: Hybrid Reasoning for Efficient Text-to-SQL with Structural Skeleton Guidance
by: Qiu, Suming, et al.
Published: (2025)
by: Qiu, Suming, et al.
Published: (2025)
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
by: Marek, Martin, et al.
Published: (2026)
by: Marek, Martin, et al.
Published: (2026)
Replay Can Provably Increase Forgetting
by: Mahdaviyeh, Yasaman, et al.
Published: (2025)
by: Mahdaviyeh, Yasaman, et al.
Published: (2025)
LocMoE: A Low-Overhead MoE for Large Language Model Training
by: Li, Jing, et al.
Published: (2024)
by: Li, Jing, et al.
Published: (2024)
ALSA: Anchors in Logit Space for Out-of-Distribution Accuracy Estimation
by: Liu, Chenzhi, et al.
Published: (2025)
by: Liu, Chenzhi, et al.
Published: (2025)
Defense without Forgetting: Continual Adversarial Defense with Anisotropic & Isotropic Pseudo Replay
by: Zhou, Yuhang, et al.
Published: (2024)
by: Zhou, Yuhang, et al.
Published: (2024)
AGGC: Adaptive Group Gradient Clipping for Stabilizing Large Language Model Training
by: Li, Zhiyuan, et al.
Published: (2026)
by: Li, Zhiyuan, et al.
Published: (2026)
CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training
by: Thede, Lukas, et al.
Published: (2026)
by: Thede, Lukas, et al.
Published: (2026)
TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale
by: Ganguli, Anurup
Published: (2026)
by: Ganguli, Anurup
Published: (2026)
A Quantitative Characterization of Forgetting in Post-Training
by: Balasubramanian, Krishnakumar, et al.
Published: (2026)
by: Balasubramanian, Krishnakumar, et al.
Published: (2026)
Spurious Forgetting in Continual Learning of Language Models
by: Zheng, Junhao, et al.
Published: (2025)
by: Zheng, Junhao, et al.
Published: (2025)
ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning
by: Jin, Chao, et al.
Published: (2026)
by: Jin, Chao, et al.
Published: (2026)
CORE: Mitigating Catastrophic Forgetting in Continual Learning through Cognitive Replay
by: Zhang, Jianshu, et al.
Published: (2024)
by: Zhang, Jianshu, et al.
Published: (2024)
SoftAdaClip: A Smooth Clipping Strategy for Fair and Private Model Training
by: Soleymani, Dorsa, et al.
Published: (2025)
by: Soleymani, Dorsa, et al.
Published: (2025)
Mapping Post-Training Forgetting in Language Models at Scale
by: Harmon, Jackson, et al.
Published: (2025)
by: Harmon, Jackson, et al.
Published: (2025)
KoReA-SFL: Knowledge Replay-based Split Federated Learning Against Catastrophic Forgetting
by: Xia, Zeke, et al.
Published: (2024)
by: Xia, Zeke, et al.
Published: (2024)
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
by: Qiu, Haiquan, et al.
Published: (2025)
by: Qiu, Haiquan, et al.
Published: (2025)
Replay-and-Forget-Free Graph Class-Incremental Learning: A Task Profiling and Prompting Approach
by: Niu, Chaoxi, et al.
Published: (2024)
by: Niu, Chaoxi, et al.
Published: (2024)
Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
by: Wang, Yuanyi, et al.
Published: (2026)
by: Wang, Yuanyi, et al.
Published: (2026)
Multi-Head Attention as a Source of Catastrophic Forgetting in MoE Transformers
by: Chen, Anrui, et al.
Published: (2026)
by: Chen, Anrui, et al.
Published: (2026)
Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training
by: Emadi, Seyed Morteza
Published: (2026)
by: Emadi, Seyed Morteza
Published: (2026)
Can LLMs Learn New Concepts Incrementally without Forgetting?
by: Zheng, Junhao, et al.
Published: (2024)
by: Zheng, Junhao, et al.
Published: (2024)
On the Plasticity and Stability for Post-Training Large Language Models
by: Qiang, Wenwen, et al.
Published: (2026)
by: Qiang, Wenwen, et al.
Published: (2026)
FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual Learning
by: Feng, Yujie, et al.
Published: (2026)
by: Feng, Yujie, et al.
Published: (2026)
MoE-DisCo:Low Economy Cost Training Mixture-of-Experts Models
by: Ye, Xin, et al.
Published: (2026)
by: Ye, Xin, et al.
Published: (2026)
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions
by: Marshall, Noah, et al.
Published: (2024)
by: Marshall, Noah, et al.
Published: (2024)
Clip-Low Increases Entropy and Clip-High Decreases Entropy in Reinforcement Learning of Large Language Models
by: Park, Jaesung R., et al.
Published: (2025)
by: Park, Jaesung R., et al.
Published: (2025)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
by: Li, Jin, et al.
Published: (2025)
by: Li, Jin, et al.
Published: (2025)
May the Forgetting Be with You: Alternate Replay for Learning with Noisy Labels
by: Millunzi, Monica, et al.
Published: (2024)
by: Millunzi, Monica, et al.
Published: (2024)
Catastrophic Forgetting Mitigation via Discrepancy-Weighted Experience Replay
by: Xu, Xinrun, et al.
Published: (2025)
by: Xu, Xinrun, et al.
Published: (2025)
MuCon: Clipped Muon Updates for LLM Training
by: Yi, Albert
Published: (2026)
by: Yi, Albert
Published: (2026)
Efficient RL Training for LLMs with Experience Replay
by: Arnal, Charles, et al.
Published: (2026)
by: Arnal, Charles, et al.
Published: (2026)
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
by: Xu, Yinggan, et al.
Published: (2026)
by: Xu, Yinggan, et al.
Published: (2026)
Stabilizing Policy Optimization via Logits Convexity
by: Chen, Hongzhan, et al.
Published: (2026)
by: Chen, Hongzhan, et al.
Published: (2026)
Regularized Gradient Clipping Provably Trains Wide and Deep Neural Networks
by: Tucat, Matteo, et al.
Published: (2024)
by: Tucat, Matteo, et al.
Published: (2024)
Prioritized Replay for RL Post-training
by: Fatemi, Mehdi
Published: (2026)
by: Fatemi, Mehdi
Published: (2026)
Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
by: Lai, Song, et al.
Published: (2025)
by: Lai, Song, et al.
Published: (2025)
Similar Items
-
Data Trajectory Alignment for LLM Domain Adaptation: A Two-Phase Synthesis Framework for Telecommunications Mathematics
by: Zhou, Zhicheng, et al.
Published: (2025) -
Agentic-KGR: Co-evolutionary Knowledge Graph Construction through Multi-Agent Reinforcement Learning
by: Li, Jing, et al.
Published: (2025) -
DeepJSONEval: Benchmarking Complex Nested JSON Data Mining for Large Language Models
by: Zhou, Zhicheng, et al.
Published: (2025) -
HES-SQL: Hybrid Reasoning for Efficient Text-to-SQL with Structural Skeleton Guidance
by: Qiu, Suming, et al.
Published: (2025) -
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
by: Marek, Martin, et al.
Published: (2026)