Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
Fuente:
arXiv
Saved in:
| Main Authors: | Glentis, Athanasios, Li, Jiaxiang, Han, Andi, Hong, Mingyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
by: Yau, Chung-Yiu, et al.
Published: (2026)
by: Yau, Chung-Yiu, et al.
Published: (2026)
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
A Minimalist Bayesian Framework for Stochastic Optimization
by: Wang, Kaizheng
Published: (2025)
by: Wang, Kaizheng
Published: (2025)
MINTS: Minimalist Thompson Sampling
by: Wang, Kaizheng
Published: (2026)
by: Wang, Kaizheng
Published: (2026)
Demystifying Manifold Constraints in LLM Pre-training
by: An, Kang, et al.
Published: (2026)
by: An, Kang, et al.
Published: (2026)
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less
by: Liu, Yuxing, et al.
Published: (2026)
by: Liu, Yuxing, et al.
Published: (2026)
From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling
by: Lin, Jianghao, et al.
Published: (2026)
by: Lin, Jianghao, et al.
Published: (2026)
Dynamic Memory Based Adaptive Optimization
by: Szegedy, Balázs, et al.
Published: (2024)
by: Szegedy, Balázs, et al.
Published: (2024)
How Memory in Optimization Algorithms Implicitly Modifies the Loss
by: Cattaneo, Matias D., et al.
Published: (2025)
by: Cattaneo, Matias D., et al.
Published: (2025)
A Queueing-Theoretic Framework for Stability Analysis of LLM Inference with KV Cache Memory Constraints
by: Nie, Chengyi, et al.
Published: (2026)
by: Nie, Chengyi, et al.
Published: (2026)
LLM Serving Optimization with Variable Prefill and Decode Lengths
by: Wang, Meixuan, et al.
Published: (2025)
by: Wang, Meixuan, et al.
Published: (2025)
Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
by: Chen, Zixi, et al.
Published: (2025)
by: Chen, Zixi, et al.
Published: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026)
by: Meterez, Alexandru, et al.
Published: (2026)
EMC$^2$: Efficient MCMC Negative Sampling for Contrastive Learning with Global Convergence
by: Yau, Chung-Yiu, et al.
Published: (2024)
by: Yau, Chung-Yiu, et al.
Published: (2024)
ECPv2: Fast, Efficient, and Scalable Global Optimization of Lipschitz Functions
by: Fourati, Fares, et al.
Published: (2025)
by: Fourati, Fares, et al.
Published: (2025)
Intersectional Fairness via Mixed-Integer Optimization
by: Němeček, Jiří, et al.
Published: (2026)
by: Němeček, Jiří, et al.
Published: (2026)
OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain Optimization Models with LLM Agents
by: Ao, Ruicheng, et al.
Published: (2026)
by: Ao, Ruicheng, et al.
Published: (2026)
Efficient Optimization with Orthogonality Constraint: a Randomized Riemannian Submanifold Method
by: Han, Andi, et al.
Published: (2025)
by: Han, Andi, et al.
Published: (2025)
Online Submodular Maximization via Online Convex Optimization
by: Salem, Tareq Si, et al.
Published: (2023)
by: Salem, Tareq Si, et al.
Published: (2023)
SPAP: Structured Pruning via Alternating Optimization and Penalty Methods
by: Hu, Hanyu, et al.
Published: (2025)
by: Hu, Hanyu, et al.
Published: (2025)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
by: Pan, Rui, et al.
Published: (2024)
by: Pan, Rui, et al.
Published: (2024)
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
by: Lau, Tim Tsz-Kit, et al.
Published: (2026)
by: Lau, Tim Tsz-Kit, et al.
Published: (2026)
Hidden Convexity of Fair PCA and Fast Solver via Eigenvalue Optimization
by: Shen, Junhui, et al.
Published: (2025)
by: Shen, Junhui, et al.
Published: (2025)
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
by: Du, Zhehang, et al.
Published: (2026)
by: Du, Zhehang, et al.
Published: (2026)
Policy Optimization in Hybrid Discrete-Continuous Action Spaces via Mixed Gradients
by: Alvo, Matias, et al.
Published: (2026)
by: Alvo, Matias, et al.
Published: (2026)
Unsupervised Training of Diffusion Models for Feasible Solution Generation in Neural Combinatorial Optimization
by: Hong, Seong-Hyun, et al.
Published: (2024)
by: Hong, Seong-Hyun, et al.
Published: (2024)
PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
Learning Concave Bid Shading Strategies in Online Auctions via Measure-valued Proximal Optimization
by: Nodozi, Iman, et al.
Published: (2025)
by: Nodozi, Iman, et al.
Published: (2025)
Beyond Minimax Rates in Group Distributionally Robust Optimization via a Novel Notion of Sparsity
by: Nguyen, Quan, et al.
Published: (2024)
by: Nguyen, Quan, et al.
Published: (2024)
Designing Service Systems from Textual Evidence
by: Ao, Ruicheng, et al.
Published: (2026)
by: Ao, Ruicheng, et al.
Published: (2026)
Parameter-Efficient Distributional RL via Normalizing Flows and a Geometry-Aware Cramér Surrogate
by: C., Simo Alami, et al.
Published: (2025)
by: C., Simo Alami, et al.
Published: (2025)
Solving Integrated Process Planning and Scheduling Problem via Graph Neural Network Based Deep Reinforcement Learning
by: Li, Hongpei, et al.
Published: (2024)
by: Li, Hongpei, et al.
Published: (2024)
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
by: Ao, Ruicheng, et al.
Published: (2025)
by: Ao, Ruicheng, et al.
Published: (2025)
Closing the Loop: Coordinating Inventory and Recommendation via Deep Reinforcement Learning on Multiple Timescales
by: Jiang, Jinyang, et al.
Published: (2025)
by: Jiang, Jinyang, et al.
Published: (2025)
Zeroth-Order Optimization Finds Flat Minima
by: Zhang, Liang, et al.
Published: (2025)
by: Zhang, Liang, et al.
Published: (2025)
Constructing Industrial-Scale Optimization Modeling Benchmark
by: Li, Zhong, et al.
Published: (2026)
by: Li, Zhong, et al.
Published: (2026)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
by: Fernando, Heshan, et al.
Published: (2024)
by: Fernando, Heshan, et al.
Published: (2024)
ControlAgent: Automating Control System Design via Novel Integration of LLM Agents and Domain Expertise
by: Guo, Xingang, et al.
Published: (2024)
by: Guo, Xingang, et al.
Published: (2024)
Optimizing the Optimizer for Physics-Informed Neural Networks and Kolmogorov-Arnold Networks
by: Kiyani, Elham, et al.
Published: (2025)
by: Kiyani, Elham, et al.
Published: (2025)
Similar Items
-
Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking
by: Glentis, Athanasios, et al.
Published: (2025) -
EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
by: Yau, Chung-Yiu, et al.
Published: (2026) -
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025) -
A Minimalist Bayesian Framework for Stochastic Optimization
by: Wang, Kaizheng
Published: (2025) -
MINTS: Minimalist Thompson Sampling
by: Wang, Kaizheng
Published: (2026)