A New Perspective on Shampoo's Preconditioner
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Morwani, Depen, Shapira, Itai, Vyas, Nikhil, Malach, Eran, Kakade, Sham, Janson, Lucas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SOAP: Improving and Stabilizing Shampoo using Adam
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024)
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024)
How Does Critical Batch Size Scale in Pre-training?
von: Zhang, Hanlin, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2024)
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026)
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
von: Abreu, Natalie, et al.
Veröffentlicht: (2025)
von: Abreu, Natalie, et al.
Veröffentlicht: (2025)
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
von: Morwani, Depen, et al.
Veröffentlicht: (2025)
von: Morwani, Depen, et al.
Veröffentlicht: (2025)
Deconstructing What Makes a Good Optimizer for Language Models
von: Zhao, Rosie, et al.
Veröffentlicht: (2024)
von: Zhao, Rosie, et al.
Veröffentlicht: (2024)
Beyond Implicit Bias: The Insignificance of SGD Noise in Online Learning
von: Vyas, Nikhil, et al.
Veröffentlicht: (2023)
von: Vyas, Nikhil, et al.
Veröffentlicht: (2023)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
von: Liu, Bingbin, et al.
Veröffentlicht: (2025)
von: Liu, Bingbin, et al.
Veröffentlicht: (2025)
Matching the Statistical Query Lower Bound for $k$-Sparse Parity Problems with Sign Stochastic Gradient Descent
von: Kou, Yiwen, et al.
Veröffentlicht: (2024)
von: Kou, Yiwen, et al.
Veröffentlicht: (2024)
LOTION: Smoothing the Optimization Landscape for Quantized Training
von: Kwun, Mujin, et al.
Veröffentlicht: (2025)
von: Kwun, Mujin, et al.
Veröffentlicht: (2025)
Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided Preconditioning
von: Li, Huan, et al.
Veröffentlicht: (2026)
von: Li, Huan, et al.
Veröffentlicht: (2026)
Online Min-Max Optimization: From Individual Regrets to Cumulative Saddle Points
von: Vyas, Abhijeet, et al.
Veröffentlicht: (2026)
von: Vyas, Abhijeet, et al.
Veröffentlicht: (2026)
The Recurrent Transformer: Greater Effective Depth and Efficient Decoding
von: Oncescu, Costin-Andrei, et al.
Veröffentlicht: (2026)
von: Oncescu, Costin-Andrei, et al.
Veröffentlicht: (2026)
Error Feedback Can Accurately Compress Preconditioners
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2023)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2023)
Feature emergence via margin maximization: case studies in algebraic tasks
von: Morwani, Depen, et al.
Veröffentlicht: (2023)
von: Morwani, Depen, et al.
Veröffentlicht: (2023)
Universal Length Generalization with Turing Programs
von: Hou, Kaiying, et al.
Veröffentlicht: (2024)
von: Hou, Kaiying, et al.
Veröffentlicht: (2024)
Repeat After Me: Transformers are Better than State Space Models at Copying
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
Efficient Graph Laplacian Estimation by Proximal Newton
von: Medvedovsky, Yakov, et al.
Veröffentlicht: (2023)
von: Medvedovsky, Yakov, et al.
Veröffentlicht: (2023)
Accelerated Parameter-Free Stochastic Optimization
von: Kreisler, Itai, et al.
Veröffentlicht: (2024)
von: Kreisler, Itai, et al.
Veröffentlicht: (2024)
A Control Theoretic Framework for Adaptive Gradient Optimizers in Machine Learning
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2022)
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2022)
New Perspectives on the Polyak Stepsize: Surrogate Functions and Negative Results
von: Orabona, Francesco, et al.
Veröffentlicht: (2025)
von: Orabona, Francesco, et al.
Veröffentlicht: (2025)
Robustness of Iteratively Pre-Conditioned Gradient-Descent Method: The Case of Distributed Linear Regression Problem
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2021)
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2021)
Iterative Pre-Conditioning for Expediting the Gradient-Descent Method: The Distributed Linear Least-Squares Problem
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2020)
von: Chakrabarti, Kushal, et al.
Veröffentlicht: (2020)
Mixture of Parrots: Experts improve memorization more than reasoning
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
von: Jelassi, Samy, et al.
Veröffentlicht: (2024)
A Pontryagin Perspective on Reinforcement Learning
von: Eberhard, Onno, et al.
Veröffentlicht: (2024)
von: Eberhard, Onno, et al.
Veröffentlicht: (2024)
A New Perspective On Denoising Based On Optimal Transport
von: Trillos, Nicolas Garcia, et al.
Veröffentlicht: (2023)
von: Trillos, Nicolas Garcia, et al.
Veröffentlicht: (2023)
The Role of Sparsity for Length Generalization in Transformers
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
Challenges in Training PINNs: A Loss Landscape Perspective
von: Rathore, Pratik, et al.
Veröffentlicht: (2024)
von: Rathore, Pratik, et al.
Veröffentlicht: (2024)
A Mirror Descent Perspective of Smoothed Sign Descent
von: Wang, Shuyang, et al.
Veröffentlicht: (2024)
von: Wang, Shuyang, et al.
Veröffentlicht: (2024)
Model-Free $μ$-Synthesis: A Nonsmooth Optimization Perspective
von: Keivan, Darioush, et al.
Veröffentlicht: (2024)
von: Keivan, Darioush, et al.
Veröffentlicht: (2024)
Langevin Dynamics: A Unified Perspective on Optimization via Lyapunov Potentials
von: Chen, August Y., et al.
Veröffentlicht: (2024)
von: Chen, August Y., et al.
Veröffentlicht: (2024)
Decentralized Bilevel Optimization: A Perspective from Transient Iteration Complexity
von: Kong, Boao, et al.
Veröffentlicht: (2024)
von: Kong, Boao, et al.
Veröffentlicht: (2024)
From Distributional Robustness to Robust Statistics: A Confidence Sets Perspective
von: Chan, Gabriel, et al.
Veröffentlicht: (2024)
von: Chan, Gabriel, et al.
Veröffentlicht: (2024)
Level Set Teleportation: An Optimization Perspective
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)
von: Mishkin, Aaron, et al.
Veröffentlicht: (2024)
Why Do We Need Warm-up? A Theoretical Perspective
von: Alimisis, Foivos, et al.
Veröffentlicht: (2025)
von: Alimisis, Foivos, et al.
Veröffentlicht: (2025)
Bounds on Perfect Node Classification: A Convex Graph Clustering Perspective
von: Shahriari-Mehr, Firooz, et al.
Veröffentlicht: (2025)
von: Shahriari-Mehr, Firooz, et al.
Veröffentlicht: (2025)
On the Stability of Nonlinear Receding Horizon Control: A Geometric Perspective
von: Westenbroek, Tyler, et al.
Veröffentlicht: (2021)
von: Westenbroek, Tyler, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
SOAP: Improving and Stabilizing Shampoo using Adam
von: Vyas, Nikhil, et al.
Veröffentlicht: (2024) -
How Does Critical Batch Size Scale in Pre-training?
von: Zhang, Hanlin, et al.
Veröffentlicht: (2024) -
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025) -
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
von: Meterez, Alexandru, et al.
Veröffentlicht: (2026) -
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
von: Meterez, Alexandru, et al.
Veröffentlicht: (2025)