Multiclass Loss Geometry Matters for Generalization of Gradient Descent in Separable Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Schliserman, Matan, Koren, Tomer |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Dimension Strikes Back with Gradients: Generalization of Gradient Methods in Stochastic Convex Optimization
by: Schliserman, Matan, et al.
Published: (2024)
by: Schliserman, Matan, et al.
Published: (2024)
Complexity of Vector-valued Prediction: From Linear Models to Stochastic Convex Optimization
by: Schliserman, Matan, et al.
Published: (2024)
by: Schliserman, Matan, et al.
Published: (2024)
Flat Minima and Generalization: Insights from Stochastic Convex Optimization
by: Schliserman, Matan, et al.
Published: (2025)
by: Schliserman, Matan, et al.
Published: (2025)
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
by: Attia, Amit, et al.
Published: (2025)
by: Attia, Amit, et al.
Published: (2025)
Optimal Rates in Continual Linear Regression via Increasing Regularization
by: Levinstein, Ran, et al.
Published: (2025)
by: Levinstein, Ran, et al.
Published: (2025)
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
by: Evron, Itay, et al.
Published: (2025)
by: Evron, Itay, et al.
Published: (2025)
The Real Price of Bandit Information in Multiclass Classification
by: Erez, Liad, et al.
Published: (2024)
by: Erez, Liad, et al.
Published: (2024)
Fast Rates for Bandit PAC Multiclass Classification
by: Erez, Liad, et al.
Published: (2024)
by: Erez, Liad, et al.
Published: (2024)
The Implicit Bias of Gradient Descent on Separable Multiclass Data
by: Ravi, Hrithik, et al.
Published: (2024)
by: Ravi, Hrithik, et al.
Published: (2024)
Rapid Overfitting of Multi-Pass Stochastic Gradient Descent in Stochastic Convex Optimization
by: Vansover-Hager, Shira, et al.
Published: (2025)
by: Vansover-Hager, Shira, et al.
Published: (2025)
Sample Complexity of Agnostic Multiclass Classification: Natarajan Dimension Strikes Back
by: Cohen, Alon, et al.
Published: (2025)
by: Cohen, Alon, et al.
Published: (2025)
Convergence of Policy Mirror Descent Beyond Compatible Function Approximation
by: Sherman, Uri, et al.
Published: (2025)
by: Sherman, Uri, et al.
Published: (2025)
The Hidden Cost of Approximation in Online Mirror Descent
by: Schlisselberg, Ofir, et al.
Published: (2025)
by: Schlisselberg, Ofir, et al.
Published: (2025)
From Contextual Combinatorial Semi-Bandits to Bandit List Classification: Improved Sample Complexity with Sparse Rewards
by: Erez, Liad, et al.
Published: (2025)
by: Erez, Liad, et al.
Published: (2025)
The Sample Complexity of Multiclass and Sparse Contextual Bandits
by: Erez, Liad, et al.
Published: (2026)
by: Erez, Liad, et al.
Published: (2026)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
Optimal Learning from Label Proportions with General Loss Functions
by: Applebaum, Lorne, et al.
Published: (2025)
by: Applebaum, Lorne, et al.
Published: (2025)
A General Reduction for High-Probability Analysis with General Light-Tailed Distributions
by: Attia, Amit, et al.
Published: (2024)
by: Attia, Amit, et al.
Published: (2024)
Any-stepsize Gradient Descent for Separable Data under Fenchel-Young Losses
by: Bao, Han, et al.
Published: (2025)
by: Bao, Han, et al.
Published: (2025)
Lower Bounds on Adversarial Robustness for Multiclass Classification with General Loss Functions
by: Trillos, Camilo Andrés García, et al.
Published: (2025)
by: Trillos, Camilo Andrés García, et al.
Published: (2025)
Learning Rate Annealing Improves Tuning Robustness in Stochastic Optimization
by: Attia, Amit, et al.
Published: (2025)
by: Attia, Amit, et al.
Published: (2025)
How Free is Parameter-Free Stochastic Optimization?
by: Attia, Amit, et al.
Published: (2024)
by: Attia, Amit, et al.
Published: (2024)
In-context Learning and Gradient Descent Revisited
by: Deutch, Gilad, et al.
Published: (2023)
by: Deutch, Gilad, et al.
Published: (2023)
The Implicit Bias of Gradient Descent on Separable Data
by: Soudry, Daniel, et al.
Published: (2017)
by: Soudry, Daniel, et al.
Published: (2017)
Online Structured Prediction with Fenchel--Young Losses and Improved Surrogate Regret for Online Multiclass Classification with Logistic Loss
by: Sakaue, Shinsaku, et al.
Published: (2024)
by: Sakaue, Shinsaku, et al.
Published: (2024)
Optimal Rates for Generalization of Gradient Descent for Deep ReLU Classification
by: Li, Yuanfan, et al.
Published: (2025)
by: Li, Yuanfan, et al.
Published: (2025)
A Theoretical Analysis of Noise Geometry in Stochastic Gradient Descent
by: Wang, Mingze, et al.
Published: (2023)
by: Wang, Mingze, et al.
Published: (2023)
Convergence and Sample Complexity of First-Order Methods for Agnostic Reinforcement Learning
by: Sherman, Uri, et al.
Published: (2025)
by: Sherman, Uri, et al.
Published: (2025)
Faster Stochastic Optimization with Arbitrary Delays via Asynchronous Mini-Batching
by: Attia, Amit, et al.
Published: (2024)
by: Attia, Amit, et al.
Published: (2024)
Multiplicative Reweighting for Robust Neural Network Optimization
by: Bar, Noga, et al.
Published: (2021)
by: Bar, Noga, et al.
Published: (2021)
Precise Asymptotic Generalization for Multiclass Classification with Overparameterized Linear Models
by: Wu, David X., et al.
Published: (2023)
by: Wu, David X., et al.
Published: (2023)
Learnable Loss Geometries with Mirror Descent for Scalable and Convergent Meta-Learning
by: Zhang, Yilang, et al.
Published: (2025)
by: Zhang, Yilang, et al.
Published: (2025)
The Multiclass Score-Oriented Loss (MultiSOL) on the Simplex
by: Marchetti, Francesco, et al.
Published: (2025)
by: Marchetti, Francesco, et al.
Published: (2025)
Large Stepsize Gradient Descent for Logistic Loss: Non-Monotonicity of the Loss Improves Optimization Efficiency
by: Wu, Jingfeng, et al.
Published: (2024)
by: Wu, Jingfeng, et al.
Published: (2024)
Type-II Saddles and Probabilistic Stability of Stochastic Gradient Descent
by: Ziyin, Liu, et al.
Published: (2023)
by: Ziyin, Liu, et al.
Published: (2023)
Exponential Convergence of (Stochastic) Gradient Descent for Separable Logistic Regression
by: Kale, Sacchit, et al.
Published: (2026)
by: Kale, Sacchit, et al.
Published: (2026)
On the Generalization of Stochastic Gradient Descent with Momentum
by: Ramezani-Kebrya, Ali, et al.
Published: (2018)
by: Ramezani-Kebrya, Ali, et al.
Published: (2018)
Rate-Optimal Policy Optimization for Linear Markov Decision Processes
by: Sherman, Uri, et al.
Published: (2023)
by: Sherman, Uri, et al.
Published: (2023)
Meta-Learning with Versatile Loss Geometries for Fast Adaptation Using Mirror Descent
by: Zhang, Yilang, et al.
Published: (2023)
by: Zhang, Yilang, et al.
Published: (2023)
Beyond Bandit Feedback in Online Multiclass Classification
by: van der Hoeven, Dirk, et al.
Published: (2021)
by: van der Hoeven, Dirk, et al.
Published: (2021)
Similar Items
-
The Dimension Strikes Back with Gradients: Generalization of Gradient Methods in Stochastic Convex Optimization
by: Schliserman, Matan, et al.
Published: (2024) -
Complexity of Vector-valued Prediction: From Linear Models to Stochastic Convex Optimization
by: Schliserman, Matan, et al.
Published: (2024) -
Flat Minima and Generalization: Insights from Stochastic Convex Optimization
by: Schliserman, Matan, et al.
Published: (2025) -
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
by: Attia, Amit, et al.
Published: (2025) -
Optimal Rates in Continual Linear Regression via Increasing Regularization
by: Levinstein, Ran, et al.
Published: (2025)