Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
Fuente:
arXiv
Saved in:
| Main Authors: | Eschenhagen, Runa, Defazio, Aaron, Lee, Tsung-Hsien, Turner, Richard E., Shi, Hao-Jun Michael |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
by: Eschenhagen, Runa, et al.
Published: (2026)
by: Eschenhagen, Runa, et al.
Published: (2026)
SOAP: Improving and Stabilizing Shampoo using Adam
by: Vyas, Nikhil, et al.
Published: (2024)
by: Vyas, Nikhil, et al.
Published: (2024)
Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization
by: Lin, Wu, et al.
Published: (2025)
by: Lin, Wu, et al.
Published: (2025)
A New Perspective on Shampoo's Preconditioner
by: Morwani, Depen, et al.
Published: (2024)
by: Morwani, Depen, et al.
Published: (2024)
4-bit Shampoo for Memory-Efficient Network Training
by: Wang, Sike, et al.
Published: (2024)
by: Wang, Sike, et al.
Published: (2024)
Pro-KLShampoo: Projected KL-Shampoo with Whitening Recovered by Orthogonalization
by: Sun, Ruotong, et al.
Published: (2026)
by: Sun, Ruotong, et al.
Published: (2026)
Influence Functions for Scalable Data Attribution in Diffusion Models
by: Mlodozeniec, Bruno, et al.
Published: (2024)
by: Mlodozeniec, Bruno, et al.
Published: (2024)
ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models
by: Defazio, Aaron
Published: (2026)
by: Defazio, Aaron
Published: (2026)
Why Gradients Rapidly Increase Near the End of Training
by: Defazio, Aaron
Published: (2025)
by: Defazio, Aaron
Published: (2025)
FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-Oriented Error for Shampoo
by: Nam, Kyunghun, et al.
Published: (2026)
by: Nam, Kyunghun, et al.
Published: (2026)
Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
by: Defazio, Aaron, et al.
Published: (2025)
by: Defazio, Aaron, et al.
Published: (2025)
Prodigy: An Expeditiously Adaptive Parameter-Free Learner
by: Mishchenko, Konstantin, et al.
Published: (2023)
by: Mishchenko, Konstantin, et al.
Published: (2023)
Optimal Linear Decay Learning Rate Schedules and Further Refinements
by: Defazio, Aaron, et al.
Published: (2023)
by: Defazio, Aaron, et al.
Published: (2023)
Better Hessians Matter: Studying the Impact of Curvature Approximations in Influence Functions
by: Hong, Steve, et al.
Published: (2025)
by: Hong, Steve, et al.
Published: (2025)
Reparametrizing Shampoo and SOAP for Subspace Basis Updates and BFloat16 Storage
by: Milligan, Alan, et al.
Published: (2026)
by: Milligan, Alan, et al.
Published: (2026)
Kronecker-Factored Approximate Curvature for Modern Neural Network Architectures
by: Eschenhagen, Runa, et al.
Published: (2023)
by: Eschenhagen, Runa, et al.
Published: (2023)
An Introduction to Transformers
by: Turner, Richard E.
Published: (2023)
by: Turner, Richard E.
Published: (2023)
Fairness Begins with State: Purifying Latent Preferences for Hierarchical Reinforcement Learning in Interactive Recommendation
by: Lu, Yun, et al.
Published: (2026)
by: Lu, Yun, et al.
Published: (2026)
Generative modeling of Sparse Approximate Inverse Preconditioners
by: Li, Mou, et al.
Published: (2024)
by: Li, Mou, et al.
Published: (2024)
TimeMixer: Decomposable Multiscale Mixing for Time Series Forecasting
by: Wang, Shiyu, et al.
Published: (2024)
by: Wang, Shiyu, et al.
Published: (2024)
DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
The Road Less Scheduled
by: Defazio, Aaron, et al.
Published: (2024)
by: Defazio, Aaron, et al.
Published: (2024)
CoLafier: Collaborative Noisy Label Purifier With Local Intrinsic Dimensionality Guidance
by: Zhang, Dongyu, et al.
Published: (2024)
by: Zhang, Dongyu, et al.
Published: (2024)
DATA: Decomposed Attention-based Task Adaptation for Rehearsal-Free Continual Learning
by: Liao, Huanxuan, et al.
Published: (2025)
by: Liao, Huanxuan, et al.
Published: (2025)
Dynamic Low-rank Approximation of Full-Matrix Preconditioner for Training Generalized Linear Models
by: Matveeva, Tatyana, et al.
Published: (2025)
by: Matveeva, Tatyana, et al.
Published: (2025)
Task-Specific Preconditioner for Cross-Domain Few-Shot Learning
by: Kang, Suhyun, et al.
Published: (2024)
by: Kang, Suhyun, et al.
Published: (2024)
Vanishing Bias Heuristic-guided Reinforcement Learning Algorithm
by: Li, Qinru, et al.
Published: (2023)
by: Li, Qinru, et al.
Published: (2023)
NeuraLSP: An Efficient and Rigorous Neural Left Singular Subspace Preconditioner for Conjugate Gradient Methods
by: Benanti, Alexander, et al.
Published: (2026)
by: Benanti, Alexander, et al.
Published: (2026)
Going Beyond Heuristics by Imposing Policy Improvement as a Constraint
by: Lee, Chi-Chang, et al.
Published: (2025)
by: Lee, Chi-Chang, et al.
Published: (2025)
Generalizable Heuristic Generation Through LLMs with Meta-Optimization
by: Shi, Yiding, et al.
Published: (2025)
by: Shi, Yiding, et al.
Published: (2025)
Spectral-factorized Positive-definite Curvature Learning for NN Training
by: Lin, Wu, et al.
Published: (2025)
by: Lin, Wu, et al.
Published: (2025)
Taming Preconditioner Drift: Unlocking the Potential of Second-Order Optimizers for Federated Learning on Non-IID Data
by: Liu, Junkang, et al.
Published: (2026)
by: Liu, Junkang, et al.
Published: (2026)
Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective
by: Lin, Wu, et al.
Published: (2024)
by: Lin, Wu, et al.
Published: (2024)
PRISM: Purified Representation and Integrated Semantic Modeling for Generative Sequential Recommendation
by: Fang, Dengzhao, et al.
Published: (2026)
by: Fang, Dengzhao, et al.
Published: (2026)
Decomposing Epistemic Uncertainty for Causal Decision Making
by: Rahman, Md Musfiqur, et al.
Published: (2026)
by: Rahman, Md Musfiqur, et al.
Published: (2026)
SDQ: Sparse Decomposed Quantization for LLM Inference
by: Jeong, Geonhwa, et al.
Published: (2024)
by: Jeong, Geonhwa, et al.
Published: (2024)
Decomposing and Editing Predictions by Modeling Model Computation
by: Shah, Harshay, et al.
Published: (2024)
by: Shah, Harshay, et al.
Published: (2024)
Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning
by: Huang, Xinting, et al.
Published: (2025)
by: Huang, Xinting, et al.
Published: (2025)
ConformaDecompose: Explaining Uncertainty via Calibration Localization
by: Yapicioglu, Fatima Rabia, et al.
Published: (2026)
by: Yapicioglu, Fatima Rabia, et al.
Published: (2026)
Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean Reference
by: Li, Jianwei, et al.
Published: (2026)
by: Li, Jianwei, et al.
Published: (2026)
Similar Items
-
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
by: Eschenhagen, Runa, et al.
Published: (2026) -
SOAP: Improving and Stabilizing Shampoo using Adam
by: Vyas, Nikhil, et al.
Published: (2024) -
Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization
by: Lin, Wu, et al.
Published: (2025) -
A New Perspective on Shampoo's Preconditioner
by: Morwani, Depen, et al.
Published: (2024) -
4-bit Shampoo for Memory-Efficient Network Training
by: Wang, Sike, et al.
Published: (2024)