DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Modoranu, Ionut-Vlad, Zmushko, Philip, Schultheis, Erik, Safaryan, Mher, Alistarh, Dan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026)
The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information
von: Wu, Diyuan, et al.
Veröffentlicht: (2024)
von: Wu, Diyuan, et al.
Veröffentlicht: (2024)
LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
von: Robert, Thomas, et al.
Veröffentlicht: (2024)
von: Robert, Thomas, et al.
Veröffentlicht: (2024)
FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of Large Language Models
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2025)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2025)
Unified Scaling Laws for Compressed Representations
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable Convergence
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2024)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2024)
LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
von: Jovanović, Andrej, et al.
Veröffentlicht: (2026)
von: Jovanović, Andrej, et al.
Veröffentlicht: (2026)
Towards Robust Scaling Laws for Optimizers
von: Volkova, Alexandra, et al.
Veröffentlicht: (2026)
von: Volkova, Alexandra, et al.
Veröffentlicht: (2026)
Error Feedback Can Accurately Compress Preconditioners
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2023)
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2023)
LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
von: Schultheis, Erik, et al.
Veröffentlicht: (2025)
von: Schultheis, Erik, et al.
Veröffentlicht: (2025)
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
von: Tabesh, Soroush, et al.
Veröffentlicht: (2025)
von: Tabesh, Soroush, et al.
Veröffentlicht: (2025)
Quartet II: Accurate LLM Pre-Training in NVFP4 by Improved Unbiased Gradient Estimation
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
Optimizers Qualitatively Alter Solutions And We Should Leverage This
von: Pascanu, Razvan, et al.
Veröffentlicht: (2025)
von: Pascanu, Razvan, et al.
Veröffentlicht: (2025)
GradSkip: Communication-Accelerated Local Gradient Methods with Better Computational Complexity
von: Maranjyan, Artavazd, et al.
Veröffentlicht: (2022)
von: Maranjyan, Artavazd, et al.
Veröffentlicht: (2022)
Grid Games: The Power of Multiple Grids for Quantizing Large Language Models
von: Egiazarian, Vage, et al.
Veröffentlicht: (2026)
von: Egiazarian, Vage, et al.
Veröffentlicht: (2026)
On Biased Compression for Distributed Learning
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2020)
von: Beznosikov, Aleksandr, et al.
Veröffentlicht: (2020)
FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable Training
von: Zmushko, Philip, et al.
Veröffentlicht: (2024)
von: Zmushko, Philip, et al.
Veröffentlicht: (2024)
Model Compression with Exact Budget Constraints via Riemannian Manifolds
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
von: Rodionov, Gleb, et al.
Veröffentlicht: (2025)
von: Rodionov, Gleb, et al.
Veröffentlicht: (2025)
Hybrid Decentralized Optimization: Leveraging Both First- and Zeroth-Order Optimizers for Faster Convergence
von: Ansaripour, Matin, et al.
Veröffentlicht: (2022)
von: Ansaripour, Matin, et al.
Veröffentlicht: (2022)
MatGPTQ: Accurate and Efficient Post-Training Matryoshka Quantization
von: Kleinegger, Maximilian, et al.
Veröffentlicht: (2026)
von: Kleinegger, Maximilian, et al.
Veröffentlicht: (2026)
Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided Preconditioning
von: Li, Huan, et al.
Veröffentlicht: (2026)
von: Li, Huan, et al.
Veröffentlicht: (2026)
Behemoth: Benchmarking Unlearning in LLMs Using Fully Synthetic Data
von: Iofinova, Eugenia, et al.
Veröffentlicht: (2026)
von: Iofinova, Eugenia, et al.
Veröffentlicht: (2026)
Efficient Data Selection at Scale via Influence Distillation
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2025)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2025)
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2025)
von: Eschenhagen, Runa, et al.
Veröffentlicht: (2025)
Communication-Efficient Federated Learning With Data and Client Heterogeneity
von: Zakerinia, Hossein, et al.
Veröffentlicht: (2022)
von: Zakerinia, Hossein, et al.
Veröffentlicht: (2022)
RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2024)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2024)
4-bit Shampoo for Memory-Efficient Network Training
von: Wang, Sike, et al.
Veröffentlicht: (2024)
von: Wang, Sike, et al.
Veröffentlicht: (2024)
Label Privacy in Split Learning for Large Models with Parameter-Efficient Training
von: Zmushko, Philip, et al.
Veröffentlicht: (2024)
von: Zmushko, Philip, et al.
Veröffentlicht: (2024)
SGD for Variational Inference: Tackling Unbounded Variance via Preconditioning and Dynamic Batching
von: Labarrière, Hippolyte, et al.
Veröffentlicht: (2026)
von: Labarrière, Hippolyte, et al.
Veröffentlicht: (2026)
Apertus LLM Family Expansion via Distillation and Quantization
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
von: Panferov, Andrei, et al.
Veröffentlicht: (2026)
Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization
von: Lin, Wu, et al.
Veröffentlicht: (2025)
von: Lin, Wu, et al.
Veröffentlicht: (2025)
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
von: Sieberling, Oliver, et al.
Veröffentlicht: (2024)
von: Sieberling, Oliver, et al.
Veröffentlicht: (2024)
Statistically-Lossless Quantization of Large Language Models
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
von: Helcig, Michael, et al.
Veröffentlicht: (2026)
Powerset Convolutional Neural Networks
von: Wendler, Chris, et al.
Veröffentlicht: (2019)
von: Wendler, Chris, et al.
Veröffentlicht: (2019)
Solving Dense Linear Systems Faster Than via Preconditioning
von: Dereziński, Michał, et al.
Veröffentlicht: (2023)
von: Dereziński, Michał, et al.
Veröffentlicht: (2023)
IBCB: Efficient Inverse Batched Contextual Bandit for Behavioral Evolution History
von: Xu, Yi, et al.
Veröffentlicht: (2024)
von: Xu, Yi, et al.
Veröffentlicht: (2024)
Faster Stochastic Optimization with Arbitrary Delays via Asynchronous Mini-Batching
von: Attia, Amit, et al.
Veröffentlicht: (2024)
von: Attia, Amit, et al.
Veröffentlicht: (2024)
SAPPHIRE: Preconditioned Stochastic Variance Reduction for Faster Large-Scale Statistical Learning
von: Sun, Jingruo, et al.
Veröffentlicht: (2025)
von: Sun, Jingruo, et al.
Veröffentlicht: (2025)
Faster Sampling from Log-Concave Densities over Polytopes via Efficient Linear Solvers
von: Mangoubi, Oren, et al.
Veröffentlicht: (2024)
von: Mangoubi, Oren, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2026) -
The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information
von: Wu, Diyuan, et al.
Veröffentlicht: (2024) -
LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
von: Robert, Thomas, et al.
Veröffentlicht: (2024) -
FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of Large Language Models
von: Modoranu, Ionut-Vlad, et al.
Veröffentlicht: (2025) -
Unified Scaling Laws for Compressed Representations
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)