High-dimensional Limit of SGD for Diagonal Linear Networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Malaxechebarría, Begoña García, Paquette, Courtney, Fazel, Maryam, Drusvyatskiy, Dmitriy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms
von: Collins-Woodfin, Elizabeth, et al.
Veröffentlicht: (2024)
von: Collins-Woodfin, Elizabeth, et al.
Veröffentlicht: (2024)
Iteratively reweighted kernel machines efficiently learn sparse functions
von: Zhu, Libin, et al.
Veröffentlicht: (2025)
von: Zhu, Libin, et al.
Veröffentlicht: (2025)
Statistical Inference for Linear Functionals of Online SGD in High-dimensional Linear Regression
von: Agrawalla, Bhavya, et al.
Veröffentlicht: (2023)
von: Agrawalla, Bhavya, et al.
Veröffentlicht: (2023)
4+3 Phases of Compute-Optimal Neural Scaling Laws
von: Paquette, Elliot, et al.
Veröffentlicht: (2024)
von: Paquette, Elliot, et al.
Veröffentlicht: (2024)
High-dimensional scaling limits and fluctuations of online least-squares SGD with smooth covariance
von: Balasubramanian, Krishnakumar, et al.
Veröffentlicht: (2023)
von: Balasubramanian, Krishnakumar, et al.
Veröffentlicht: (2023)
Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models
von: Zhu, Libin, et al.
Veröffentlicht: (2026)
von: Zhu, Libin, et al.
Veröffentlicht: (2026)
Dimension-adapted Momentum Outscales SGD
von: Ferbach, Damien, et al.
Veröffentlicht: (2025)
von: Ferbach, Damien, et al.
Veröffentlicht: (2025)
Phases of Muon: When Muon Eclipses SignSGD
von: Paquette, Elliot, et al.
Veröffentlicht: (2026)
von: Paquette, Elliot, et al.
Veröffentlicht: (2026)
A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
von: Davis, Damek, et al.
Veröffentlicht: (2026)
von: Davis, Damek, et al.
Veröffentlicht: (2026)
When do spectral gradient updates help in deep learning?
von: Davis, Damek, et al.
Veröffentlicht: (2025)
von: Davis, Damek, et al.
Veröffentlicht: (2025)
Invariant Kernels: Rank Stabilization and Generalization Across Dimensions
von: Díaz, Mateo, et al.
Veröffentlicht: (2025)
von: Díaz, Mateo, et al.
Veröffentlicht: (2025)
A Piecewise Lyapunov Analysis of Sub-quadratic SGD: Applications to Robust and Quantile Regression
von: Zhang, Yixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Yixuan, et al.
Veröffentlicht: (2025)
The radius of statistical efficiency
von: Cutler, Joshua, et al.
Veröffentlicht: (2024)
von: Cutler, Joshua, et al.
Veröffentlicht: (2024)
Decentralized Sparse Linear Regression via Gradient-Tracking: Linear Convergence and Statistical Guarantees
von: Maros, Marie, et al.
Veröffentlicht: (2022)
von: Maros, Marie, et al.
Veröffentlicht: (2022)
Trajectory-Restricted Optimization Conditions and Geometry-Aware Linear Convergence
von: Chaudhry, Faris, et al.
Veröffentlicht: (2026)
von: Chaudhry, Faris, et al.
Veröffentlicht: (2026)
Robustly Learning Monotone Generalized Linear Models via Data Augmentation
von: Zarifis, Nikos, et al.
Veröffentlicht: (2025)
von: Zarifis, Nikos, et al.
Veröffentlicht: (2025)
A Novel Framework for Policy Mirror Descent with General Parameterization and Linear Convergence
von: Alfano, Carlo, et al.
Veröffentlicht: (2023)
von: Alfano, Carlo, et al.
Veröffentlicht: (2023)
On the Sample Complexity of Set Membership Estimation for Linear Systems with Disturbances Bounded by Convex Sets
von: Xu, Haonan, et al.
Veröffentlicht: (2024)
von: Xu, Haonan, et al.
Veröffentlicht: (2024)
Joint Learning of Linear Dynamical Systems under Smoothness Constraints
von: Tyagi, Hemant
Veröffentlicht: (2024)
von: Tyagi, Hemant
Veröffentlicht: (2024)
High dimensional analysis reveals conservative sharpening and a stochastic edge of stability
von: Agarwala, Atish, et al.
Veröffentlicht: (2024)
von: Agarwala, Atish, et al.
Veröffentlicht: (2024)
Online Experimental Design With Estimation-Regret Trade-off Under Network Interference
von: Zhang, Zhiheng, et al.
Veröffentlicht: (2024)
von: Zhang, Zhiheng, et al.
Veröffentlicht: (2024)
High-probability Convergence Bounds for Nonlinear Stochastic Gradient Descent Under Heavy-tailed Noise
von: Armacki, Aleksandar, et al.
Veröffentlicht: (2023)
von: Armacki, Aleksandar, et al.
Veröffentlicht: (2023)
Beyond Maximum Likelihood: Variational Inequality Estimation for Generalized Linear Models
von: Zhu, Linglingzhi, et al.
Veröffentlicht: (2025)
von: Zhu, Linglingzhi, et al.
Veröffentlicht: (2025)
A New Perspective On Denoising Based On Optimal Transport
von: Trillos, Nicolas Garcia, et al.
Veröffentlicht: (2023)
von: Trillos, Nicolas Garcia, et al.
Veröffentlicht: (2023)
Function Gradient Approximation with Random Shallow ReLU Networks with Control Applications
von: Lamperski, Andrew, et al.
Veröffentlicht: (2024)
von: Lamperski, Andrew, et al.
Veröffentlicht: (2024)
Markov Kernels, Distances and Optimal Control: A Parable of Linear Quadratic Non-Gaussian Distribution Steering
von: Teter, Alexis M. H., et al.
Veröffentlicht: (2025)
von: Teter, Alexis M. H., et al.
Veröffentlicht: (2025)
High-probability sample complexities for policy evaluation with linear function approximation
von: Li, Gen, et al.
Veröffentlicht: (2023)
von: Li, Gen, et al.
Veröffentlicht: (2023)
Logarithmic-time Schedules for Scaling Language Models with Momentum
von: Ferbach, Damien, et al.
Veröffentlicht: (2026)
von: Ferbach, Damien, et al.
Veröffentlicht: (2026)
A Spectral Framework for Closed-Form Relative Density Estimation
von: Bach, Francis
Veröffentlicht: (2026)
von: Bach, Francis
Veröffentlicht: (2026)
Risk reversal for least squares estimators under nested convex constraints
von: Al-Ghattas, Omar
Veröffentlicht: (2026)
von: Al-Ghattas, Omar
Veröffentlicht: (2026)
Robust stochastic first order methods in heavy-tailed noise via medoid mini-batch gradient sampling
von: Vukovic, Manojlo, et al.
Veröffentlicht: (2026)
von: Vukovic, Manojlo, et al.
Veröffentlicht: (2026)
Data-Efficient Non-Gaussian Semi-Nonparametric Density Estimation for Nonlinear Dynamical Systems
von: Liao, Aaron R., et al.
Veröffentlicht: (2026)
von: Liao, Aaron R., et al.
Veröffentlicht: (2026)
Computation of Least Trimmed Squares: A Branch-and-Bound framework with Hyperplane Arrangement Enhancements
von: Meng, Xiang, et al.
Veröffentlicht: (2026)
von: Meng, Xiang, et al.
Veröffentlicht: (2026)
Frequentist Regret Analysis of Gaussian Process Thompson Sampling via Fractional Posteriors
von: Roy, Somjit, et al.
Veröffentlicht: (2026)
von: Roy, Somjit, et al.
Veröffentlicht: (2026)
Continuous-time reinforcement learning: ellipticity enables model-free value function approximation
von: Mou, Wenlong
Veröffentlicht: (2026)
von: Mou, Wenlong
Veröffentlicht: (2026)
Robust Assortment Optimization from Observational Data
von: Lu, Miao, et al.
Veröffentlicht: (2026)
von: Lu, Miao, et al.
Veröffentlicht: (2026)
Stochastic Optimization with Optimal Importance Sampling
von: Aolaritei, Liviu, et al.
Veröffentlicht: (2025)
von: Aolaritei, Liviu, et al.
Veröffentlicht: (2025)
A Theory of Feature Learning in Kernel Models
von: Chen, Yunlu, et al.
Veröffentlicht: (2023)
von: Chen, Yunlu, et al.
Veröffentlicht: (2023)
Joint learning of a network of linear dynamical systems via total variation penalization
von: Donnat, Claire, et al.
Veröffentlicht: (2025)
von: Donnat, Claire, et al.
Veröffentlicht: (2025)
A review of NMF, PLSA, LBA, EMA, and LCA with a focus on the identifiability issue
von: Qi, Qianqian, et al.
Veröffentlicht: (2025)
von: Qi, Qianqian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms
von: Collins-Woodfin, Elizabeth, et al.
Veröffentlicht: (2024) -
Iteratively reweighted kernel machines efficiently learn sparse functions
von: Zhu, Libin, et al.
Veröffentlicht: (2025) -
Statistical Inference for Linear Functionals of Online SGD in High-dimensional Linear Regression
von: Agrawalla, Bhavya, et al.
Veröffentlicht: (2023) -
4+3 Phases of Compute-Optimal Neural Scaling Laws
von: Paquette, Elliot, et al.
Veröffentlicht: (2024) -
High-dimensional scaling limits and fluctuations of online least-squares SGD with smooth covariance
von: Balasubramanian, Krishnakumar, et al.
Veröffentlicht: (2023)