MSign: An Optimizer Preventing Training Instability in Large Language Models via Stable Rank Restoration
Fuente:
arXiv
Saved in:
| Main Authors: | Ren, Lianhai, Ding, Yucheng, Liu, Xiao, Li, Qianxiao, Cheng, Peng, Gong, Yeyun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unifying back-propagation and forward-forward algorithms through model predictive control
by: Ren, Lianhai, et al.
Published: (2024)
by: Ren, Lianhai, et al.
Published: (2024)
Optimizing Large Language Model Training Using FP4 Quantization
by: Wang, Ruizhe, et al.
Published: (2025)
by: Wang, Ruizhe, et al.
Published: (2025)
Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
by: Wang, Ruizhe, et al.
Published: (2025)
by: Wang, Ruizhe, et al.
Published: (2025)
From Generalization Analysis to Optimization Designs for State Space Models
by: Liu, Fusheng, et al.
Published: (2024)
by: Liu, Fusheng, et al.
Published: (2024)
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
by: Wang, Shida, et al.
Published: (2023)
by: Wang, Shida, et al.
Published: (2023)
Autocorrelation Matters: Understanding the Role of Initialization Schemes for State Space Models
by: Liu, Fusheng, et al.
Published: (2024)
by: Liu, Fusheng, et al.
Published: (2024)
APOLLO: An Optimized Training Approach for Long-form Numerical Reasoning
by: Sun, Jiashuo, et al.
Published: (2022)
by: Sun, Jiashuo, et al.
Published: (2022)
Approximation Rate of the Transformer Architecture for Sequence Modeling
by: Jiang, Haotian, et al.
Published: (2023)
by: Jiang, Haotian, et al.
Published: (2023)
Large Language Model Compression with Global Rank and Sparsity Optimization
by: Zhou, Changhai, et al.
Published: (2025)
by: Zhou, Changhai, et al.
Published: (2025)
BOOST: BOttleneck-Optimized Scalable Training Framework for Low-Rank Large Language Models
by: Wang, Zhengyang, et al.
Published: (2025)
by: Wang, Zhengyang, et al.
Published: (2025)
Accelerating Training of Autoregressive Video Generation Models via Local Optimization with Representation Continuity
by: Zhou, Yucheng, et al.
Published: (2026)
by: Zhou, Yucheng, et al.
Published: (2026)
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
by: Shin, Haebin, et al.
Published: (2025)
by: Shin, Haebin, et al.
Published: (2025)
Machine Unlearning under Retain-Forget Entanglement
by: Cheng, Jingpu, et al.
Published: (2026)
by: Cheng, Jingpu, et al.
Published: (2026)
Learning task-specific predictive models for scientific computing
by: Yin, Jianyuan, et al.
Published: (2025)
by: Yin, Jianyuan, et al.
Published: (2025)
Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability
by: Kim, Bum Jun, et al.
Published: (2026)
by: Kim, Bum Jun, et al.
Published: (2026)
CoRA: Optimizing Low-Rank Adaptation with Common Subspace of Large Language Models
by: Xiao, Xiaojun, et al.
Published: (2024)
by: Xiao, Xiaojun, et al.
Published: (2024)
ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity
by: Li, Jiaxi, et al.
Published: (2026)
by: Li, Jiaxi, et al.
Published: (2026)
C-LoRA: Contextual Low-Rank Adaptation for Uncertainty Estimation in Large Language Models
by: Rahmati, Amir Hossein, et al.
Published: (2025)
by: Rahmati, Amir Hossein, et al.
Published: (2025)
An Overview of Low-Rank Structures in the Training and Adaptation of Large Models
by: Balzano, Laura, et al.
Published: (2025)
by: Balzano, Laura, et al.
Published: (2025)
Restoring Pruned Large Language Models via Lost Component Compensation
by: Feng, Zijian, et al.
Published: (2025)
by: Feng, Zijian, et al.
Published: (2025)
Probabilistic Lipschitzness and the Stable Rank for Comparing Explanation Models
by: Simpson, Lachlan, et al.
Published: (2024)
by: Simpson, Lachlan, et al.
Published: (2024)
Beam Prediction based on Large Language Models
by: Sheng, Yucheng, et al.
Published: (2024)
by: Sheng, Yucheng, et al.
Published: (2024)
Can Watermarking Large Language Models Prevent Copyrighted Text Generation and Hide Training Data?
by: Panaitescu-Liess, Michael-Andrei, et al.
Published: (2024)
by: Panaitescu-Liess, Michael-Andrei, et al.
Published: (2024)
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
by: Lee, Jung Hyun, et al.
Published: (2024)
by: Lee, Jung Hyun, et al.
Published: (2024)
Boost Post-Training Quantization via Null Space Optimization for Large Language Models
by: Zhao, Jiaqi, et al.
Published: (2025)
by: Zhao, Jiaqi, et al.
Published: (2025)
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
by: Yang, Kailai, et al.
Published: (2025)
by: Yang, Kailai, et al.
Published: (2025)
Accelerating Legacy Numerical Solvers by Non-intrusive Gradient-based Meta-solving
by: Arisaka, Sohei, et al.
Published: (2024)
by: Arisaka, Sohei, et al.
Published: (2024)
Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection
by: Zhan, Zheng, et al.
Published: (2025)
by: Zhan, Zheng, et al.
Published: (2025)
Adaptive Feature-based Low-Rank Compression of Large Language Models via Bayesian Optimization
by: Ji, Yixin, et al.
Published: (2024)
by: Ji, Yixin, et al.
Published: (2024)
Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training
by: Zhou, Wenjie, et al.
Published: (2026)
by: Zhou, Wenjie, et al.
Published: (2026)
Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models
by: Wang, Liangyu, et al.
Published: (2025)
by: Wang, Liangyu, et al.
Published: (2025)
Entropy-Gated Selective Policy Optimization:Token-Level Gradient Allocation for Hybrid Training of Large Language Models
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
LayerNorm Induces Recency Bias in Transformer Decoders
by: Kim, Junu, et al.
Published: (2025)
by: Kim, Junu, et al.
Published: (2025)
The Effect of Depth on the Expressivity of Deep Linear State-Space Models
by: Bao, Zeyu, et al.
Published: (2025)
by: Bao, Zeyu, et al.
Published: (2025)
D-QRELO: Training- and Data-Free Delta Compression for Large Language Models via Quantization and Residual Low-Rank Approximation
by: Li, Junlin, et al.
Published: (2026)
by: Li, Junlin, et al.
Published: (2026)
Numerical Investigation of Sequence Modeling Theory using Controllable Memory Functions
by: Jiang, Haotian, et al.
Published: (2025)
by: Jiang, Haotian, et al.
Published: (2025)
A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models
by: Chen, Yiming, et al.
Published: (2025)
by: Chen, Yiming, et al.
Published: (2025)
Training-Free Bayesianization for Low-Rank Adapters of Large Language Models
by: Shi, Haizhou, et al.
Published: (2024)
by: Shi, Haizhou, et al.
Published: (2024)
Dynamic Low-Rank Sparse Adaptation for Large Language Models
by: Huang, Weizhong, et al.
Published: (2025)
by: Huang, Weizhong, et al.
Published: (2025)
Group Causal Policy Optimization for Post-Training Large Language Models
by: Gu, Ziyin, et al.
Published: (2025)
by: Gu, Ziyin, et al.
Published: (2025)
Similar Items
-
Unifying back-propagation and forward-forward algorithms through model predictive control
by: Ren, Lianhai, et al.
Published: (2024) -
Optimizing Large Language Model Training Using FP4 Quantization
by: Wang, Ruizhe, et al.
Published: (2025) -
Beyond Sunk Costs: Boosting LLM Pre-training Efficiency via Orthogonal Growth of Mixture-of-Experts
by: Wang, Ruizhe, et al.
Published: (2025) -
From Generalization Analysis to Optimization Designs for State Space Models
by: Liu, Fusheng, et al.
Published: (2024) -
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization
by: Wang, Shida, et al.
Published: (2023)