SOAP: Improving and Stabilizing Shampoo using Adam
Fuente:
arXiv
Saved in:
| Main Authors: | Vyas, Nikhil, Morwani, Depen, Zhao, Rosie, Kwun, Mujin, Shapira, Itai, Brandfonbrener, David, Janson, Lucas, Kakade, Sham |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deconstructing What Makes a Good Optimizer for Language Models
by: Zhao, Rosie, et al.
Published: (2024)
by: Zhao, Rosie, et al.
Published: (2024)
A New Perspective on Shampoo's Preconditioner
by: Morwani, Depen, et al.
Published: (2024)
by: Morwani, Depen, et al.
Published: (2024)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
by: Morwani, Depen, et al.
Published: (2025)
by: Morwani, Depen, et al.
Published: (2025)
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
by: Abreu, Natalie, et al.
Published: (2025)
by: Abreu, Natalie, et al.
Published: (2025)
LOTION: Smoothing the Optimization Landscape for Quantized Training
by: Kwun, Mujin, et al.
Published: (2025)
by: Kwun, Mujin, et al.
Published: (2025)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
Beyond Implicit Bias: The Insignificance of SGD Noise in Online Learning
by: Vyas, Nikhil, et al.
Published: (2023)
by: Vyas, Nikhil, et al.
Published: (2023)
How Does Critical Batch Size Scale in Pre-training?
by: Zhang, Hanlin, et al.
Published: (2024)
by: Zhang, Hanlin, et al.
Published: (2024)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
by: Oncescu, Costin-Andrei, et al.
Published: (2026)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
by: Liu, Bingbin, et al.
Published: (2025)
by: Liu, Bingbin, et al.
Published: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026)
by: Meterez, Alexandru, et al.
Published: (2026)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Feature emergence via margin maximization: case studies in algebraic tasks
by: Morwani, Depen, et al.
Published: (2023)
by: Morwani, Depen, et al.
Published: (2023)
Characterization and Mitigation of Training Instabilities in Microscaling Formats
by: Su, Huangyuan, et al.
Published: (2025)
by: Su, Huangyuan, et al.
Published: (2025)
CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training
by: Brandfonbrener, David, et al.
Published: (2024)
by: Brandfonbrener, David, et al.
Published: (2024)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2025)
by: Meterez, Alexandru, et al.
Published: (2025)
Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
by: Jelassi, Samy, et al.
Published: (2026)
by: Jelassi, Samy, et al.
Published: (2026)
GQ-VAE: A gated quantized VAE for learning variable length tokens
by: Datta, Theo, et al.
Published: (2025)
by: Datta, Theo, et al.
Published: (2025)
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
by: Eschenhagen, Runa, et al.
Published: (2025)
by: Eschenhagen, Runa, et al.
Published: (2025)
Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization
by: Lin, Wu, et al.
Published: (2025)
by: Lin, Wu, et al.
Published: (2025)
Learning Hidden Markov Models Using Conditional Samples
by: Kakade, Sham M., et al.
Published: (2023)
by: Kakade, Sham M., et al.
Published: (2023)
Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
by: Oncescu, Costin-Andrei, et al.
Published: (2024)
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
by: Zhang, Hanlin, et al.
Published: (2026)
by: Zhang, Hanlin, et al.
Published: (2026)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
by: Jin, Jikai, et al.
Published: (2025)
by: Jin, Jikai, et al.
Published: (2025)
Pairwise Calibrated Rewards for Pluralistic Alignment
by: Halpern, Daniel, et al.
Published: (2025)
by: Halpern, Daniel, et al.
Published: (2025)
Soup to go: mitigating forgetting during continual learning with model averaging
by: Kleiman, Anat, et al.
Published: (2025)
by: Kleiman, Anat, et al.
Published: (2025)
Universal Length Generalization with Turing Programs
by: Hou, Kaiying, et al.
Published: (2024)
by: Hou, Kaiying, et al.
Published: (2024)
Scaling Laws for Imitation Learning in Single-Agent Games
by: Tuyls, Jens, et al.
Published: (2023)
by: Tuyls, Jens, et al.
Published: (2023)
Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes
by: Kamal, Sadia, et al.
Published: (2025)
by: Kamal, Sadia, et al.
Published: (2025)
Mixture of Parrots: Experts improve memorization more than reasoning
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
4-bit Shampoo for Memory-Efficient Network Training
by: Wang, Sike, et al.
Published: (2024)
by: Wang, Sike, et al.
Published: (2024)
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
by: Eschenhagen, Runa, et al.
Published: (2026)
by: Eschenhagen, Runa, et al.
Published: (2026)
Pro-KLShampoo: Projected KL-Shampoo with Whitening Recovered by Orthogonalization
by: Sun, Ruotong, et al.
Published: (2026)
by: Sun, Ruotong, et al.
Published: (2026)
Cognitive models can reveal interpretable value trade-offs in language models
by: Murthy, Sonia K., et al.
Published: (2025)
by: Murthy, Sonia K., et al.
Published: (2025)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
by: Qi, Zhenting, et al.
Published: (2024)
by: Qi, Zhenting, et al.
Published: (2024)
SOAP-RL: Sequential Option Advantage Propagation for Reinforcement Learning in POMDP Environments
by: Ishida, Shu, et al.
Published: (2024)
by: Ishida, Shu, et al.
Published: (2024)
Q-Probe: A Lightweight Approach to Reward Maximization for Language Models
by: Li, Kenneth, et al.
Published: (2024)
by: Li, Kenneth, et al.
Published: (2024)
Similar Items
-
Deconstructing What Makes a Good Optimizer for Language Models
by: Zhao, Rosie, et al.
Published: (2024) -
A New Perspective on Shampoo's Preconditioner
by: Morwani, Depen, et al.
Published: (2024) -
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
by: Morwani, Depen, et al.
Published: (2025) -
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
by: Abreu, Natalie, et al.
Published: (2025) -
LOTION: Smoothing the Optimization Landscape for Quantized Training
by: Kwun, Mujin, et al.
Published: (2025)