Accelerating LLM Pre-Training through Flat-Direction Dynamics Enhancement
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Shuchen, Hu, Rizhen, Wang, Mingze, Sun, Mou, Wang, Xue, Yuan, Kun, Wen, Zaiwen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grouter: Decoupling Routing from Representation for Accelerated MoE Training
by: Xu, Yuqi, et al.
Published: (2026)
by: Xu, Yuqi, et al.
Published: (2026)
Synergistic Intra- and Cross-Layer Regularization Losses for MoE Expert Specialization
by: Hu, Rizhen, et al.
Published: (2026)
by: Hu, Rizhen, et al.
Published: (2026)
A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models
by: Chen, Yiming, et al.
Published: (2025)
by: Chen, Yiming, et al.
Published: (2025)
OptProver: Bridging Olympiad and Optimization through Continual Training in Formal Theorem Proving
by: Li, Chenyi, et al.
Published: (2026)
by: Li, Chenyi, et al.
Published: (2026)
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
Accelerating Neural Network Training Along Sharp and Flat Directions
by: Zakarin, Daniyar, et al.
Published: (2025)
by: Zakarin, Daniyar, et al.
Published: (2025)
Accelerating Optimization via Differentiable Stopping Time
by: Xie, Zhonglin, et al.
Published: (2025)
by: Xie, Zhonglin, et al.
Published: (2025)
MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
by: Hu, Rizhen, et al.
Published: (2025)
by: Hu, Rizhen, et al.
Published: (2025)
Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures
by: Chen, Yiming, et al.
Published: (2024)
by: Chen, Yiming, et al.
Published: (2024)
Spatiotemporal Graph Learning with Direct Volumetric Information Passing and Feature Enhancement
by: Mi, Yuan, et al.
Published: (2024)
by: Mi, Yuan, et al.
Published: (2024)
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models
by: Wang, Mingze, et al.
Published: (2026)
by: Wang, Mingze, et al.
Published: (2026)
FlatQuant: Flatness Matters for LLM Quantization
by: Sun, Yuxuan, et al.
Published: (2024)
by: Sun, Yuxuan, et al.
Published: (2024)
Mixture-of-Channels: Exploiting Sparse FFNs for Efficient LLMs Pre-Training and Inference
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
Noise Consistency Training: A Native Approach for One-Step Generator in Learning Additional Controls
by: Luo, Yihong, et al.
Published: (2025)
by: Luo, Yihong, et al.
Published: (2025)
GradPower: Powering Gradients for Faster Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)
by: Wang, Jinbo, et al.
Published: (2025)
SPARKLE: A Unified Single-Loop Primal-Dual Framework for Decentralized Bilevel Optimization
by: Zhu, Shuchen, et al.
Published: (2024)
by: Zhu, Shuchen, et al.
Published: (2024)
Decentralized Bilevel Optimization: A Perspective from Transient Iteration Complexity
by: Kong, Boao, et al.
Published: (2024)
by: Kong, Boao, et al.
Published: (2024)
Combatting Dimensional Collapse in LLM Pre-Training Data via Diversified File Selection
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
LLM4TS: Aligning Pre-Trained LLMs as Data-Efficient Time-Series Forecasters
by: Chang, Ching, et al.
Published: (2023)
by: Chang, Ching, et al.
Published: (2023)
Accelerated Natural Gradient Method for Parametric Manifold Optimization
by: Li, Chenyi, et al.
Published: (2025)
by: Li, Chenyi, et al.
Published: (2025)
To 2:4 Sparsity and Beyond: Neuron-level Activation Function to Accelerate LLM Pre-Training
by: Madhyastha, Meghana, et al.
Published: (2026)
by: Madhyastha, Meghana, et al.
Published: (2026)
Training Instabilities Induce Flatness Bias in Gradient Descent
by: Wang, Lawrence, et al.
Published: (2025)
by: Wang, Lawrence, et al.
Published: (2025)
Gauss-Newton Temporal Difference Learning with Nonlinear Function Approximation
by: Ke, Zhifa, et al.
Published: (2023)
by: Ke, Zhifa, et al.
Published: (2023)
Meta-Prompt Optimization for LLM-Based Sequential Decision Making
by: Kong, Mingze, et al.
Published: (2025)
by: Kong, Mingze, et al.
Published: (2025)
Accelerating Transformer Pre-training with 2:4 Sparsity
by: Hu, Yuezhou, et al.
Published: (2024)
by: Hu, Yuezhou, et al.
Published: (2024)
From Dionysius Emerges Apollo -- Learning Patterns and Abstractions from Perceptual Sequences
by: Wu, Shuchen
Published: (2025)
by: Wu, Shuchen
Published: (2025)
On the Trade-off between Flatness and Optimization in Distributed Learning
by: Cao, Ying, et al.
Published: (2024)
by: Cao, Ying, et al.
Published: (2024)
Investigating the Pre-Training Dynamics of In-Context Learning: Task Recognition vs. Task Learning
by: Wang, Xiaolei, et al.
Published: (2024)
by: Wang, Xiaolei, et al.
Published: (2024)
GWT: Scalable Optimizer State Compression for Large Language Model Training
by: Wen, Ziqing, et al.
Published: (2025)
by: Wen, Ziqing, et al.
Published: (2025)
Accelerating Diffusion Sampling with Optimized Time Steps
by: Xue, Shuchen, et al.
Published: (2024)
by: Xue, Shuchen, et al.
Published: (2024)
On the Expressive Power of Mixture-of-Experts for Structured Complex Tasks
by: Wang, Mingze, et al.
Published: (2025)
by: Wang, Mingze, et al.
Published: (2025)
A Theoretical Analysis of Noise Geometry in Stochastic Gradient Descent
by: Wang, Mingze, et al.
Published: (2023)
by: Wang, Mingze, et al.
Published: (2023)
Understanding the Expressive Power and Mechanisms of Transformer for Sequence Modeling
by: Wang, Mingze, et al.
Published: (2024)
by: Wang, Mingze, et al.
Published: (2024)
Enhancing Pre-Trained Model-Based Class-Incremental Learning through Neural Collapse
by: He, Kun, et al.
Published: (2025)
by: He, Kun, et al.
Published: (2025)
An Improved Finite-time Analysis of Temporal Difference Learning with Deep Neural Networks
by: Ke, Zhifa, et al.
Published: (2024)
by: Ke, Zhifa, et al.
Published: (2024)
Sharpness-Aware Minimization Efficiently Selects Flatter Minima Late in Training
by: Zhou, Zhanpeng, et al.
Published: (2024)
by: Zhou, Zhanpeng, et al.
Published: (2024)
Unlocking Out-of-Distribution Generalization in Dynamics through Physics-Guided Augmentation
by: Xu, Fan, et al.
Published: (2025)
by: Xu, Fan, et al.
Published: (2025)
FOAM: Blocked State Folding for Memory-Efficient LLM Training
by: Wen, Ziqing, et al.
Published: (2025)
by: Wen, Ziqing, et al.
Published: (2025)
Unicron: Economizing Self-Healing LLM Training at Scale
by: He, Tao, et al.
Published: (2023)
by: He, Tao, et al.
Published: (2023)
Convergence Analysis of Stochastic Gradient Descent with MCMC Estimators
by: Li, Tianyou, et al.
Published: (2023)
by: Li, Tianyou, et al.
Published: (2023)
Similar Items
-
Grouter: Decoupling Routing from Representation for Accelerated MoE Training
by: Xu, Yuqi, et al.
Published: (2026) -
Synergistic Intra- and Cross-Layer Regularization Losses for MoE Expert Specialization
by: Hu, Rizhen, et al.
Published: (2026) -
A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models
by: Chen, Yiming, et al.
Published: (2025) -
OptProver: Bridging Olympiad and Optimization through Continual Training in Formal Theorem Proving
by: Li, Chenyi, et al.
Published: (2026) -
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
by: Wang, Jinbo, et al.
Published: (2025)