The Implicit Bias of Gradient Descent on Separable Data
Fuente:
arXiv
Saved in:
| Main Authors: | Soudry, Daniel, Hoffer, Elad, Nacson, Mor Shpigel, Gunasekar, Suriya, Srebro, Nathan |
|---|---|
| Format: | Preprint |
| Published: |
2017
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers
by: Buzaglo, Gon, et al.
Published: (2024)
by: Buzaglo, Gon, et al.
Published: (2024)
The Implicit Bias of Gradient Descent on Separable Multiclass Data
by: Ravi, Hrithik, et al.
Published: (2024)
by: Ravi, Hrithik, et al.
Published: (2024)
The Price of Implicit Bias in Adversarially Robust Generalization
by: Tsilivis, Nikolaos, et al.
Published: (2024)
by: Tsilivis, Nikolaos, et al.
Published: (2024)
DocVLM: Make Your VLM an Efficient Reader
by: Nacson, Mor Shpigel, et al.
Published: (2024)
by: Nacson, Mor Shpigel, et al.
Published: (2024)
Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
by: Chmiel, Brian, et al.
Published: (2021)
by: Chmiel, Brian, et al.
Published: (2021)
Temperature is All You Need for Generalization in Langevin Dynamics and other Markov Processes
by: Harel, Itamar, et al.
Published: (2025)
by: Harel, Itamar, et al.
Published: (2025)
Retrieval from Within: An Intrinsic Capability of Attention-Based Models
by: Hoffer, Elad, et al.
Published: (2026)
by: Hoffer, Elad, et al.
Published: (2026)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
Research Program: Theory of Learning in Dynamical Systems
by: Hazan, Elad, et al.
Published: (2025)
by: Hazan, Elad, et al.
Published: (2025)
The Implicit Bias of Steepest Descent with Mini-batch Stochastic Gradient
by: Li, Jichu, et al.
Published: (2026)
by: Li, Jichu, et al.
Published: (2026)
Provable Tempered Overfitting of Minimal Nets and Typical Nets
by: Harel, Itamar, et al.
Published: (2024)
by: Harel, Itamar, et al.
Published: (2024)
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
by: Evron, Itay, et al.
Published: (2025)
by: Evron, Itay, et al.
Published: (2025)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
by: Cai, Yuhang, et al.
Published: (2025)
by: Cai, Yuhang, et al.
Published: (2025)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
by: Jung, Hyunji, et al.
Published: (2025)
by: Jung, Hyunji, et al.
Published: (2025)
The Implicit Bias of Adam on Separable Data
by: Zhang, Chenyang, et al.
Published: (2024)
by: Zhang, Chenyang, et al.
Published: (2024)
On the Complexity of Learning Sparse Functions with Statistical and Gradient Queries
by: Joshi, Nirmit, et al.
Published: (2024)
by: Joshi, Nirmit, et al.
Published: (2024)
PLUMAGE: Probabilistic Low rank Unbiased Min Variance Gradient Estimator for Efficient Large Model Training
by: Haroush, Matan, et al.
Published: (2025)
by: Haroush, Matan, et al.
Published: (2025)
Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks
by: Li, Binghui, et al.
Published: (2024)
by: Li, Binghui, et al.
Published: (2024)
Quantifying Overfitting along the Regularization Path for Two-Part-Code MDL in Supervised Classification
by: Zhu, Xiaohan, et al.
Published: (2025)
by: Zhu, Xiaohan, et al.
Published: (2025)
Implicit Bias of Mirror Flow on Separable Data
by: Pesme, Scott, et al.
Published: (2024)
by: Pesme, Scott, et al.
Published: (2024)
Depth Separation in Norm-Bounded Infinite-Width Neural Networks
by: Parkinson, Suzanna, et al.
Published: (2024)
by: Parkinson, Suzanna, et al.
Published: (2024)
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
by: Meng, Si Yi, et al.
Published: (2024)
by: Meng, Si Yi, et al.
Published: (2024)
Shift is Good: Mismatched Data Mixing Improves Test Performance
by: Medvedev, Marko, et al.
Published: (2025)
by: Medvedev, Marko, et al.
Published: (2025)
Any-stepsize Gradient Descent for Separable Data under Fenchel-Young Losses
by: Bao, Han, et al.
Published: (2025)
by: Bao, Han, et al.
Published: (2025)
Implicit Bias and Convergence of Matrix Stochastic Mirror Descent
by: Akhtiamov, Danil, et al.
Published: (2026)
by: Akhtiamov, Danil, et al.
Published: (2026)
Workspace Optimization: How to Train Your Agent
by: Sarafian, Elad, et al.
Published: (2026)
by: Sarafian, Elad, et al.
Published: (2026)
Noisy Interpolation Learning with Shallow Univariate ReLU Networks
by: Joshi, Nirmit, et al.
Published: (2023)
by: Joshi, Nirmit, et al.
Published: (2023)
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
by: Medvedev, Marko, et al.
Published: (2024)
by: Medvedev, Marko, et al.
Published: (2024)
Training Instabilities Induce Flatness Bias in Gradient Descent
by: Wang, Lawrence, et al.
Published: (2025)
by: Wang, Lawrence, et al.
Published: (2025)
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
by: Tsilivis, Nikolaos, et al.
Published: (2024)
by: Tsilivis, Nikolaos, et al.
Published: (2024)
Towards The Implicit Bias on Multiclass Separable Data Under Norm Constraints
by: Xie, Shengping, et al.
Published: (2026)
by: Xie, Shengping, et al.
Published: (2026)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
by: Chmiel, Brian, et al.
Published: (2022)
by: Chmiel, Brian, et al.
Published: (2022)
How Does the ReLU Activation Affect the Implicit Bias of Gradient Descent on High-dimensional Neural Network Regression?
by: Lai, Kuo-Wei, et al.
Published: (2026)
by: Lai, Kuo-Wei, et al.
Published: (2026)
Overfitting and Generalizing with (PAC) Bayesian Prediction in Noisy Binary Classification
by: Zhu, Xiaohan, et al.
Published: (2026)
by: Zhu, Xiaohan, et al.
Published: (2026)
Recursive Models for Long-Horizon Reasoning
by: Yang, Chenxiao, et al.
Published: (2026)
by: Yang, Chenxiao, et al.
Published: (2026)
Understanding the Implicit Regularization of Gradient Descent in Over-parameterized Models
by: Ma, Jianhao, et al.
Published: (2025)
by: Ma, Jianhao, et al.
Published: (2025)
Multiclass Loss Geometry Matters for Generalization of Gradient Descent in Separable Classification
by: Schliserman, Matan, et al.
Published: (2025)
by: Schliserman, Matan, et al.
Published: (2025)
Tight Bounds on the Binomial CDF, and the Minimum of i.i.d Binomials, in terms of KL-Divergence
by: Zhu, Xiaohan, et al.
Published: (2025)
by: Zhu, Xiaohan, et al.
Published: (2025)
Turning Stale Gradients into Stable Gradients: Coherent Coordinate Descent with Implicit Landscape Smoothing for Lightweight Zeroth-Order Optimization
by: Liang, Chen, et al.
Published: (2026)
by: Liang, Chen, et al.
Published: (2026)
Exponential Convergence of (Stochastic) Gradient Descent for Separable Logistic Regression
by: Kale, Sacchit, et al.
Published: (2026)
by: Kale, Sacchit, et al.
Published: (2026)
Similar Items
-
How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers
by: Buzaglo, Gon, et al.
Published: (2024) -
The Implicit Bias of Gradient Descent on Separable Multiclass Data
by: Ravi, Hrithik, et al.
Published: (2024) -
The Price of Implicit Bias in Adversarially Robust Generalization
by: Tsilivis, Nikolaos, et al.
Published: (2024) -
DocVLM: Make Your VLM an Efficient Reader
by: Nacson, Mor Shpigel, et al.
Published: (2024) -
Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
by: Chmiel, Brian, et al.
Published: (2021)