Navigating LLM Valley: From AdamW to Memory-Efficient and Matrix-Based Optimizers
Fuente:
arXiv
Saved in:
| Main Author: | Ranganath, Aditya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizer-Induced Mode Connectivity: From AdamW to Muon
by: Zhang, Fangzhao, et al.
Published: (2026)
by: Zhang, Fangzhao, et al.
Published: (2026)
APOLLO: SGD-like Memory, AdamW-level Performance
by: Zhu, Hanqing, et al.
Published: (2024)
by: Zhu, Hanqing, et al.
Published: (2024)
Beyond First-Order: Training LLMs with Stochastic Conjugate Subgradients and AdamW
by: Zhang, Di, et al.
Published: (2025)
by: Zhang, Di, et al.
Published: (2025)
How to set AdamW's weight decay as you scale model and dataset size
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
First-Passage Prediction of Grokking Delay: ACalibrated Law under AdamW with Causal Validation
by: Khanh, Truong Xuan, et al.
Published: (2026)
by: Khanh, Truong Xuan, et al.
Published: (2026)
DP-FedAdamW: An Efficient Optimizer for Differentially Private Federated Large Models
by: Liu, Jin, et al.
Published: (2026)
by: Liu, Jin, et al.
Published: (2026)
Uniform Scaling Limits in AdamW-Trained Transformers
by: Gibson, William, et al.
Published: (2026)
by: Gibson, William, et al.
Published: (2026)
Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization
by: Xie, Shuo, et al.
Published: (2024)
by: Xie, Shuo, et al.
Published: (2024)
FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large Models
by: Liu, Junkang, et al.
Published: (2025)
by: Liu, Junkang, et al.
Published: (2025)
HomeAdam: Adam and AdamW Algorithms Sometimes Go Home to Obtain Better Provable Generalization
by: Huang, Feihu, et al.
Published: (2026)
by: Huang, Feihu, et al.
Published: (2026)
Resource-Efficient Iterative LLM-Based NAS with Feedback Memory
by: Gu, Xiaojie, et al.
Published: (2026)
by: Gu, Xiaojie, et al.
Published: (2026)
ProtAlign: Contrastive learning paradigm for Sequence and structure alignment
by: Ranganath, Aditya, et al.
Published: (2026)
by: Ranganath, Aditya, et al.
Published: (2026)
In-Run Data Shapley for Adam Optimizer
by: Ding, Meng, et al.
Published: (2026)
by: Ding, Meng, et al.
Published: (2026)
Memory-Efficient LLM Training with Dynamic Sparsity: From Stability to Practical Scaling
by: Xiao, Qiao, et al.
Published: (2026)
by: Xiao, Qiao, et al.
Published: (2026)
Promoting Exploration in Memory-Augmented Adam using Critical Momenta
by: Malviya, Pranshu, et al.
Published: (2023)
by: Malviya, Pranshu, et al.
Published: (2023)
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments
by: Cai, Zhijie, et al.
Published: (2026)
by: Cai, Zhijie, et al.
Published: (2026)
From $O(mn)$ to $O(r^2)$: Two-Sided Low-Rank Communication for Adam in Distributed Training with Memory Efficiency
by: Dang, Sizhe, et al.
Published: (2026)
by: Dang, Sizhe, et al.
Published: (2026)
FlashOptim: Optimizers for Memory-Efficient Training
by: Ortiz, Jose Javier Gonzalez, et al.
Published: (2026)
by: Ortiz, Jose Javier Gonzalez, et al.
Published: (2026)
KV Cache Quantization for Self-Forcing Video Generation: A 33-Method Empirical Study
by: Ranganath, Suraj, et al.
Published: (2026)
by: Ranganath, Suraj, et al.
Published: (2026)
Explanations that reveal all through the definition of encoding
by: Puli, Aahlad, et al.
Published: (2024)
by: Puli, Aahlad, et al.
Published: (2024)
DP-AdamW: Investigating Decoupled Weight Decay and Bias Correction in Private Deep Learning
by: Chooi, Jay, et al.
Published: (2025)
by: Chooi, Jay, et al.
Published: (2025)
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors
by: Ranganath, Suraj, et al.
Published: (2026)
by: Ranganath, Suraj, et al.
Published: (2026)
Memory as a Markov Matrix: Sample Efficient Knowledge Expansion via Token-to-Dictionary Mapping
by: Pethkar, Kaustubh, et al.
Published: (2026)
by: Pethkar, Kaustubh, et al.
Published: (2026)
FOAM: Blocked State Folding for Memory-Efficient LLM Training
by: Wen, Ziqing, et al.
Published: (2025)
by: Wen, Ziqing, et al.
Published: (2025)
AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training
by: Zhang, Huishuai, et al.
Published: (2025)
by: Zhang, Huishuai, et al.
Published: (2025)
Muon Outperforms Adam in Tail-End Associative Memory Learning
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
LOGICPO: Efficient Translation of NL-based Logical Problems to FOL using LLMs and Preference Optimization
by: Viswanadha, Koushik, et al.
Published: (2025)
by: Viswanadha, Koushik, et al.
Published: (2025)
HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading
by: Luo, Cheng, et al.
Published: (2025)
by: Luo, Cheng, et al.
Published: (2025)
LOOKAT: Lookup-Optimized Key-Attention for Memory-Efficient Transformers
by: Karmore, Aryan
Published: (2026)
by: Karmore, Aryan
Published: (2026)
SMMF: Square-Matricized Momentum Factorization for Memory-Efficient Optimization
by: Park, Kwangryeol, et al.
Published: (2024)
by: Park, Kwangryeol, et al.
Published: (2024)
The Epochal Sawtooth Phenomenon: Unveiling Training Loss Oscillations in Adam and Other Optimizers
by: Liu, Qi, et al.
Published: (2024)
by: Liu, Qi, et al.
Published: (2024)
From Soliloquy to Agora: Memory-Enhanced LLM Agents with Decentralized Debate for Optimization Modeling
by: Lin, Jianghao, et al.
Published: (2026)
by: Lin, Jianghao, et al.
Published: (2026)
Memory-Based Advantage Shaping for LLM-Guided Reinforcement Learning
by: Nourzad, Narjes, et al.
Published: (2026)
by: Nourzad, Narjes, et al.
Published: (2026)
Exploratory Memory-Augmented LLM Agent via Hybrid On- and Off-Policy Optimization
by: Liu, Zeyuan, et al.
Published: (2026)
by: Liu, Zeyuan, et al.
Published: (2026)
Scalable Parameter and Memory Efficient Pretraining for LLM: Recent Algorithmic Advances and Benchmarking
by: Glentis, Athanasios, et al.
Published: (2025)
by: Glentis, Athanasios, et al.
Published: (2025)
RouterBench: A Benchmark for Multi-LLM Routing System
by: Hu, Qitian Jason, et al.
Published: (2024)
by: Hu, Qitian Jason, et al.
Published: (2024)
MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
by: Liu, Yuxi, et al.
Published: (2025)
by: Liu, Yuxi, et al.
Published: (2025)
Memory-Efficient Gradient Unrolling for Large-Scale Bi-level Optimization
by: Shen, Qianli, et al.
Published: (2024)
by: Shen, Qianli, et al.
Published: (2024)
A Physics-Inspired Optimizer: Velocity Regularized Adam
by: Vaidhyanathan, Pranav, et al.
Published: (2025)
by: Vaidhyanathan, Pranav, et al.
Published: (2025)
Similar Items
-
Optimizer-Induced Mode Connectivity: From AdamW to Muon
by: Zhang, Fangzhao, et al.
Published: (2026) -
APOLLO: SGD-like Memory, AdamW-level Performance
by: Zhu, Hanqing, et al.
Published: (2024) -
Beyond First-Order: Training LLMs with Stochastic Conjugate Subgradients and AdamW
by: Zhang, Di, et al.
Published: (2025) -
How to set AdamW's weight decay as you scale model and dataset size
by: Wang, Xi, et al.
Published: (2024) -
First-Passage Prediction of Grokking Delay: ACalibrated Law under AdamW with Causal Validation
by: Khanh, Truong Xuan, et al.
Published: (2026)