Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Jihwan, Song, Dogyoon, Yun, Chulhee |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
por: Tao, Hongyi, et al.
Publicado: (2026)
por: Tao, Hongyi, et al.
Publicado: (2026)
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
por: Yu, Dingzhi, et al.
Publicado: (2026)
por: Yu, Dingzhi, et al.
Publicado: (2026)
Phases of Muon: When Muon Eclipses SignSGD
por: Paquette, Elliot, et al.
Publicado: (2026)
por: Paquette, Elliot, et al.
Publicado: (2026)
Does SGD really happen in tiny subspaces?
por: Song, Minhak, et al.
Publicado: (2024)
por: Song, Minhak, et al.
Publicado: (2024)
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
por: Petrov, Egor, et al.
Publicado: (2025)
por: Petrov, Egor, et al.
Publicado: (2025)
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing
por: Wagner, Dominik, et al.
Publicado: (2024)
por: Wagner, Dominik, et al.
Publicado: (2024)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
por: Liao, Fangshuo, et al.
Publicado: (2026)
por: Liao, Fangshuo, et al.
Publicado: (2026)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
por: Baek, Beomhan, et al.
Publicado: (2025)
por: Baek, Beomhan, et al.
Publicado: (2025)
Linear attention is (maybe) all you need (to understand transformer optimization)
por: Ahn, Kwangjun, et al.
Publicado: (2023)
por: Ahn, Kwangjun, et al.
Publicado: (2023)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
por: Song, Minhak, et al.
Publicado: (2025)
por: Song, Minhak, et al.
Publicado: (2025)
A Simplified Analysis of SGD for Linear Regression with Weight Averaging
por: Meterez, Alexandru, et al.
Publicado: (2025)
por: Meterez, Alexandru, et al.
Publicado: (2025)
Sign-SGD via Parameter-Free Optimization
por: Medyakov, Daniil, et al.
Publicado: (2025)
por: Medyakov, Daniil, et al.
Publicado: (2025)
Understanding SGD with Exponential Moving Average: A Case Study in Linear Regression
por: Li, Xuheng, et al.
Publicado: (2025)
por: Li, Xuheng, et al.
Publicado: (2025)
Accelerating Single-Pass SGD for Generalized Linear Prediction
por: Chen, Qian, et al.
Publicado: (2026)
por: Chen, Qian, et al.
Publicado: (2026)
Does SGD Seek Flatness or Sharpness? An Exactly Solvable Model
por: Xu, Yizhou, et al.
Publicado: (2026)
por: Xu, Yizhou, et al.
Publicado: (2026)
Statistical Inference for Linear Functionals of Online SGD in High-dimensional Linear Regression
por: Agrawalla, Bhavya, et al.
Publicado: (2023)
por: Agrawalla, Bhavya, et al.
Publicado: (2023)
Demystifying SGD with Doubly Stochastic Gradients
por: Kim, Kyurae, et al.
Publicado: (2024)
por: Kim, Kyurae, et al.
Publicado: (2024)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
por: Dahan, Tehila, et al.
Publicado: (2023)
por: Dahan, Tehila, et al.
Publicado: (2023)
Making SGD Parameter-Free
por: Carmon, Yair, et al.
Publicado: (2022)
por: Carmon, Yair, et al.
Publicado: (2022)
On the Trajectories of SGD Without Replacement
por: Beneventano, Pierfrancesco
Publicado: (2023)
por: Beneventano, Pierfrancesco
Publicado: (2023)
SignSGD with Federated Voting
por: Park, Chanho, et al.
Publicado: (2024)
por: Park, Chanho, et al.
Publicado: (2024)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
por: Xie, Shengping, et al.
Publicado: (2025)
por: Xie, Shengping, et al.
Publicado: (2025)
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity
por: Tyurin, Alexander, et al.
Publicado: (2024)
por: Tyurin, Alexander, et al.
Publicado: (2024)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
por: Jung, Hyunji, et al.
Publicado: (2025)
por: Jung, Hyunji, et al.
Publicado: (2025)
Dimension-adapted Momentum Outscales SGD
por: Ferbach, Damien, et al.
Publicado: (2025)
por: Ferbach, Damien, et al.
Publicado: (2025)
Heavy-Tail Phenomenon in Decentralized SGD
por: Gurbuzbalaban, Mert, et al.
Publicado: (2022)
por: Gurbuzbalaban, Mert, et al.
Publicado: (2022)
High-dimensional Limit of SGD for Diagonal Linear Networks
por: Malaxechebarría, Begoña García, et al.
Publicado: (2026)
por: Malaxechebarría, Begoña García, et al.
Publicado: (2026)
Muon Outperforms Adam in Tail-End Associative Memory Learning
por: Wang, Shuche, et al.
Publicado: (2025)
por: Wang, Shuche, et al.
Publicado: (2025)
Does Worst-Performing Agent Lead the Pack? Analyzing Agent Dynamics in Unified Distributed SGD
por: Hu, Jie, et al.
Publicado: (2024)
por: Hu, Jie, et al.
Publicado: (2024)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
por: Vasudeva, Bhavya, et al.
Publicado: (2025)
por: Vasudeva, Bhavya, et al.
Publicado: (2025)
Can SGD Handle Heavy-Tailed Noise?
por: Fatkhullin, Ilyas, et al.
Publicado: (2025)
por: Fatkhullin, Ilyas, et al.
Publicado: (2025)
SGD with memory: fundamental properties and stochastic acceleration
por: Yarotsky, Dmitry, et al.
Publicado: (2024)
por: Yarotsky, Dmitry, et al.
Publicado: (2024)
Effective Frontiers: A Unification of Neural Scaling Laws
por: Zou, Jiaxuan, et al.
Publicado: (2026)
por: Zou, Jiaxuan, et al.
Publicado: (2026)
Stochastic Extragradient with Flip-Flop Shuffling & Anchoring: Provable Improvements
por: Chae, Jiseok, et al.
Publicado: (2024)
por: Chae, Jiseok, et al.
Publicado: (2024)
Incremental Gradient Descent with Small Epoch Counts is Surprisingly Slow on Ill-Conditioned Problems
por: Kim, Yujun, et al.
Publicado: (2025)
por: Kim, Yujun, et al.
Publicado: (2025)
How Does Critical Batch Size Scale in Pre-training?
por: Zhang, Hanlin, et al.
Publicado: (2024)
por: Zhang, Hanlin, et al.
Publicado: (2024)
Optimal Projection-Free Adaptive SGD for Matrix Optimization
por: Kovalev, Dmitry
Publicado: (2026)
por: Kovalev, Dmitry
Publicado: (2026)
On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization
por: Sahu, Sharan, et al.
Publicado: (2026)
por: Sahu, Sharan, et al.
Publicado: (2026)
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD
por: Andreyev, Arseniy, et al.
Publicado: (2024)
por: Andreyev, Arseniy, et al.
Publicado: (2024)
From Gradient Clipping to Normalization for Heavy Tailed SGD
por: Hübler, Florian, et al.
Publicado: (2024)
por: Hübler, Florian, et al.
Publicado: (2024)
Ejemplares similares
-
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
por: Tao, Hongyi, et al.
Publicado: (2026) -
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
por: Yu, Dingzhi, et al.
Publicado: (2026) -
Phases of Muon: When Muon Eclipses SignSGD
por: Paquette, Elliot, et al.
Publicado: (2026) -
Does SGD really happen in tiny subspaces?
por: Song, Minhak, et al.
Publicado: (2024) -
Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order
por: Petrov, Egor, et al.
Publicado: (2025)