Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters
Fuente:
arXiv
Guardado en:
| Autores principales: | Yukhimchuk, Alexander, Kolar, Mladen, Takáč, Martin, Choudhury, Sayantan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
por: Choudhury, Sayantan, et al.
Publicado: (2026)
por: Choudhury, Sayantan, et al.
Publicado: (2026)
Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity
por: Gorbunov, Eduard, et al.
Publicado: (2024)
por: Gorbunov, Eduard, et al.
Publicado: (2024)
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
por: Chezhegov, Savelii, et al.
Publicado: (2024)
por: Chezhegov, Savelii, et al.
Publicado: (2024)
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024)
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024)
Parameter-free Clipped Gradient Descent Meets Polyak
por: Takezawa, Yuki, et al.
Publicado: (2024)
por: Takezawa, Yuki, et al.
Publicado: (2024)
Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad
por: Choudhury, Sayantan, et al.
Publicado: (2024)
por: Choudhury, Sayantan, et al.
Publicado: (2024)
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024)
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024)
Extragradient Method for $(L_0, L_1)$-Lipschitz Root-finding Problems
por: Choudhury, Sayantan, et al.
Publicado: (2025)
por: Choudhury, Sayantan, et al.
Publicado: (2025)
The Ky Fan Norms and Beyond: Dual Norms and Combinations for Matrix Optimization
por: Kravatskiy, Alexey, et al.
Publicado: (2025)
por: Kravatskiy, Alexey, et al.
Publicado: (2025)
Communication-Efficient Gradient Descent-Accent Methods for Distributed Variational Inequalities: Unified Analysis and Local Updates
por: Zhang, Siqi, et al.
Publicado: (2023)
por: Zhang, Siqi, et al.
Publicado: (2023)
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
por: Riabinin, Artem, et al.
Publicado: (2026)
por: Riabinin, Artem, et al.
Publicado: (2026)
Preconditioned Norms: A Unified Framework for Steepest Descent, Quasi-Newton and Adaptive Methods
por: Veprikov, Andrey, et al.
Publicado: (2025)
por: Veprikov, Andrey, et al.
Publicado: (2025)
Fully Stochastic Trust-Region Sequential Quadratic Programming for Equality-Constrained Optimization Problems
por: Fang, Yuchen, et al.
Publicado: (2022)
por: Fang, Yuchen, et al.
Publicado: (2022)
Random-reshuffled SARAH does not need a full gradient computations
por: Beznosikov, Aleksandr, et al.
Publicado: (2021)
por: Beznosikov, Aleksandr, et al.
Publicado: (2021)
From Gradient Clipping to Normalization for Heavy Tailed SGD
por: Hübler, Florian, et al.
Publicado: (2024)
por: Hübler, Florian, et al.
Publicado: (2024)
Multiplayer Federated Learning: Reaching Equilibrium with Less Communication
por: Yoon, TaeHo, et al.
Publicado: (2025)
por: Yoon, TaeHo, et al.
Publicado: (2025)
Stochastic Gradient Methods with Preconditioned Updates
por: Sadiev, Abdurakhmon, et al.
Publicado: (2022)
por: Sadiev, Abdurakhmon, et al.
Publicado: (2022)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024)
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024)
Regularized Gradient Clipping Provably Trains Wide and Deep Neural Networks
por: Tucat, Matteo, et al.
Publicado: (2024)
por: Tucat, Matteo, et al.
Publicado: (2024)
Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning
por: Zhang, Dake, et al.
Publicado: (2024)
por: Zhang, Dake, et al.
Publicado: (2024)
Muon Optimizes Under Spectral Norm Constraints
por: Chen, Lizhang, et al.
Publicado: (2025)
por: Chen, Lizhang, et al.
Publicado: (2025)
Trust-Region Sequential Quadratic Programming for Stochastic Optimization with Random Models
por: Fang, Yuchen, et al.
Publicado: (2024)
por: Fang, Yuchen, et al.
Publicado: (2024)
Clipped Gradient Methods for Nonsmooth Convex Optimization under Heavy-Tailed Noise: A Refined Analysis
por: Liu, Zijian
Publicado: (2025)
por: Liu, Zijian
Publicado: (2025)
Small Gradient Norm Regret for Online Convex Optimization
por: Gao, Wenzhi, et al.
Publicado: (2026)
por: Gao, Wenzhi, et al.
Publicado: (2026)
Nonconvex Stochastic Optimization under Heavy-Tailed Noises: Optimal Convergence without Gradient Clipping
por: Liu, Zijian, et al.
Publicado: (2024)
por: Liu, Zijian, et al.
Publicado: (2024)
Differentially Private Clipped-SGD: High-Probability Convergence with Arbitrary Clipping Level
por: Khah, Saleh Vatan, et al.
Publicado: (2025)
por: Khah, Saleh Vatan, et al.
Publicado: (2025)
Revisiting Gradient Normalization and Clipping for Nonconvex SGD under Heavy-Tailed Noise: Necessity, Sufficiency, and Acceleration
por: Sun, Tao, et al.
Publicado: (2024)
por: Sun, Tao, et al.
Publicado: (2024)
Federated Learning Can Find Friends That Are Advantageous
por: Tupitsa, Nazarii, et al.
Publicado: (2024)
por: Tupitsa, Nazarii, et al.
Publicado: (2024)
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size
por: Ostroukhov, Petr, et al.
Publicado: (2024)
por: Ostroukhov, Petr, et al.
Publicado: (2024)
Simple Stepsize for Quasi-Newton Methods with Global Convergence Guarantees
por: Agafonov, Artem, et al.
Publicado: (2025)
por: Agafonov, Artem, et al.
Publicado: (2025)
Convergence of Alternating Gradient Descent for Matrix Factorization
por: Ward, Rachel, et al.
Publicado: (2023)
por: Ward, Rachel, et al.
Publicado: (2023)
SANIA: Polyak-type Optimization Framework Leads to Scale Invariant Stochastic Algorithms
por: Abdukhakimov, Farshed, et al.
Publicado: (2023)
por: Abdukhakimov, Farshed, et al.
Publicado: (2023)
LoFT: Low-Rank Adaptation That Behaves Like Full Fine-Tuning
por: Tastan, Nurbek, et al.
Publicado: (2025)
por: Tastan, Nurbek, et al.
Publicado: (2025)
Non-Asymptotic Global Convergence of PPO-Clip
por: Liu, Yin, et al.
Publicado: (2025)
por: Liu, Yin, et al.
Publicado: (2025)
Bayesian Optimization with Structured Measurements: A Vector-Valued RKHS Framework
por: Wang, Wenbin, et al.
Publicado: (2026)
por: Wang, Wenbin, et al.
Publicado: (2026)
Robust and Fast Training via Per-Sample Clipping
por: Nobile, Davide, et al.
Publicado: (2026)
por: Nobile, Davide, et al.
Publicado: (2026)
Exploring New Frontiers in Vertical Federated Learning: the Role of Saddle Point Reformulation
por: Beznosikov, Aleksandr, et al.
Publicado: (2026)
por: Beznosikov, Aleksandr, et al.
Publicado: (2026)
Preconditioned Gradient Descent for Over-Parameterized Nonconvex Matrix Factorization
por: Zhang, Gavin, et al.
Publicado: (2025)
por: Zhang, Gavin, et al.
Publicado: (2025)
Convergence of Gradient Descent with Small Initialization for Unregularized Matrix Completion
por: Ma, Jianhao, et al.
Publicado: (2024)
por: Ma, Jianhao, et al.
Publicado: (2024)
Gradient-Free Approaches is a Key to an Efficient Interaction with Markovian Stochasticity
por: Prokhorov, Boris, et al.
Publicado: (2026)
por: Prokhorov, Boris, et al.
Publicado: (2026)
Ejemplares similares
-
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
por: Choudhury, Sayantan, et al.
Publicado: (2026) -
Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity
por: Gorbunov, Eduard, et al.
Publicado: (2024) -
Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed
por: Chezhegov, Savelii, et al.
Publicado: (2024) -
AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods
por: Lau, Tim Tsz-Kit, et al.
Publicado: (2024) -
Parameter-free Clipped Gradient Descent Meets Polyak
por: Takezawa, Yuki, et al.
Publicado: (2024)