Simplicity Bias of Two-Layer Networks beyond Linearly Separable Data
Fuente:
arXiv
Saved in:
| Main Authors: | Tsoy, Nikita, Konstantinov, Nikola |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Measuring Localization of Shortcuts in Deep Networks
by: Tsoy, Nikita, et al.
Published: (2025)
by: Tsoy, Nikita, et al.
Published: (2025)
Implicit Bias of Mirror Flow on Separable Data
by: Pesme, Scott, et al.
Published: (2024)
by: Pesme, Scott, et al.
Published: (2024)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
Towards The Implicit Bias on Multiclass Separable Data Under Norm Constraints
by: Xie, Shengping, et al.
Published: (2026)
by: Xie, Shengping, et al.
Published: (2026)
Hidden Minima in Two-Layer ReLU Networks
by: Arjevani, Yossi
Published: (2023)
by: Arjevani, Yossi
Published: (2023)
Convex Formulations for Training Two-Layer ReLU Neural Networks
by: Prakhya, Karthik, et al.
Published: (2024)
by: Prakhya, Karthik, et al.
Published: (2024)
Bias and Extrapolation in Markovian Linear Stochastic Approximation with Constant Stepsizes
by: Huo, Dongyan, et al.
Published: (2022)
by: Huo, Dongyan, et al.
Published: (2022)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
by: Jung, Hyunji, et al.
Published: (2025)
by: Jung, Hyunji, et al.
Published: (2025)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
by: Xu, Xianliang, et al.
Published: (2024)
by: Xu, Xianliang, et al.
Published: (2024)
Mean-Field Limits for Two-Layer Neural Networks Trained with Consensus-Based Optimization
by: De Deyn, William, et al.
Published: (2025)
by: De Deyn, William, et al.
Published: (2025)
Criteria and Bias of Parameterized Linear Regression under Edge of Stability Regime
by: Zhang, Peiyuan, et al.
Published: (2024)
by: Zhang, Peiyuan, et al.
Published: (2024)
Adversarial Training of Two-Layer Polynomial and ReLU Activation Networks via Convex Optimization
by: Kuelbs, Daniel, et al.
Published: (2024)
by: Kuelbs, Daniel, et al.
Published: (2024)
Homotopy Relaxation Training Algorithms for Infinite-Width Two-Layer ReLU Neural Networks
by: Yang, Yahong, et al.
Published: (2023)
by: Yang, Yahong, et al.
Published: (2023)
Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
by: Min, Hancheng, et al.
Published: (2025)
by: Min, Hancheng, et al.
Published: (2025)
Large Stepsize Gradient Descent for Non-Homogeneous Two-Layer Networks: Margin Improvement and Fast Optimization
by: Cai, Yuhang, et al.
Published: (2024)
by: Cai, Yuhang, et al.
Published: (2024)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
by: Baek, Beomhan, et al.
Published: (2025)
by: Baek, Beomhan, et al.
Published: (2025)
Global Convergence of SGD On Two Layer Neural Nets
by: Gopalani, Pulkit, et al.
Published: (2022)
by: Gopalani, Pulkit, et al.
Published: (2022)
On the Impact of Performative Risk Minimization for Binary Random Variables
by: Tsoy, Nikita, et al.
Published: (2025)
by: Tsoy, Nikita, et al.
Published: (2025)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
by: Cai, Yuhang, et al.
Published: (2025)
by: Cai, Yuhang, et al.
Published: (2025)
Global Convergence of SGD For Logistic Loss on Two Layer Neural Nets
by: Gopalani, Pulkit, et al.
Published: (2023)
by: Gopalani, Pulkit, et al.
Published: (2023)
Optimizer's Information Criterion: Dissecting and Correcting Bias in Data-Driven Optimization
by: Iyengar, Garud, et al.
Published: (2023)
by: Iyengar, Garud, et al.
Published: (2023)
Optimization Insights into Deep Diagonal Linear Networks
by: Labarrière, Hippolyte, et al.
Published: (2024)
by: Labarrière, Hippolyte, et al.
Published: (2024)
Two-Timescale Optimization Framework for Sparse-Feedback Linear-Quadratic Optimal Control
by: Feng, Lechen, et al.
Published: (2024)
by: Feng, Lechen, et al.
Published: (2024)
Symplectic Inductive Bias for Data-Driven Target Reachability in Hamiltonian Systems
by: Ouyang, Zhuo, et al.
Published: (2026)
by: Ouyang, Zhuo, et al.
Published: (2026)
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
by: Meng, Si Yi, et al.
Published: (2024)
by: Meng, Si Yi, et al.
Published: (2024)
Accelerated Methods with Complexity Separation Under Data Similarity for Federated Learning Problems
by: Bylinkin, Dmitry, et al.
Published: (2026)
by: Bylinkin, Dmitry, et al.
Published: (2026)
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
by: Beneventano, Pierfrancesco, et al.
Published: (2025)
by: Beneventano, Pierfrancesco, et al.
Published: (2025)
Shuffling the Data, Stretching the Step-size: Sharper Bias in constant step-size SGD
by: Emmanouilidis, Konstantinos, et al.
Published: (2026)
by: Emmanouilidis, Konstantinos, et al.
Published: (2026)
SensLI: Sensitivity-Based Layer Insertion for Neural Networks
by: Kreis, Leonie, et al.
Published: (2023)
by: Kreis, Leonie, et al.
Published: (2023)
Two-Timescale Linear Stochastic Approximation: Constant Stepsizes Go a Long Way
by: Kwon, Jeongyeol, et al.
Published: (2024)
by: Kwon, Jeongyeol, et al.
Published: (2024)
Communication Efficient Federated Learning with Linear Convergence on Heterogeneous Data
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
Physics-Informed Neural Networks with Hard Linear Equality Constraints
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Convergence Analysis for Learning Orthonormal Deep Linear Neural Networks
by: Qin, Zhen, et al.
Published: (2023)
by: Qin, Zhen, et al.
Published: (2023)
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
by: Das, Mohua, et al.
Published: (2026)
by: Das, Mohua, et al.
Published: (2026)
Tight Finite Time Bounds of Two-Time-Scale Linear Stochastic Approximation with Markovian Noise
by: Haque, Shaan Ul, et al.
Published: (2023)
by: Haque, Shaan Ul, et al.
Published: (2023)
Data-Driven Adversarial Online Control for Unknown Linear Systems
by: Liu, Zishun, et al.
Published: (2023)
by: Liu, Zishun, et al.
Published: (2023)
Interpolation Conditions for Data Consistency and Prediction in Noisy Linear Systems
by: Vanelli, Martina, et al.
Published: (2025)
by: Vanelli, Martina, et al.
Published: (2025)
What Data Enables Optimal Decisions? An Exact Characterization for Linear Optimization
by: Bennouna, Omar, et al.
Published: (2025)
by: Bennouna, Omar, et al.
Published: (2025)
A Nonlinear Separation Principle via Contraction Theory: Applications to Neural Networks, Control, and Learning
by: Gokhale, Anand, et al.
Published: (2026)
by: Gokhale, Anand, et al.
Published: (2026)
On subdifferential chain rule of matrix factorization and beyond
by: Guan, Jiewen, et al.
Published: (2024)
by: Guan, Jiewen, et al.
Published: (2024)
Similar Items
-
On Measuring Localization of Shortcuts in Deep Networks
by: Tsoy, Nikita, et al.
Published: (2025) -
Implicit Bias of Mirror Flow on Separable Data
by: Pesme, Scott, et al.
Published: (2024) -
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025) -
Towards The Implicit Bias on Multiclass Separable Data Under Norm Constraints
by: Xie, Shengping, et al.
Published: (2026) -
Hidden Minima in Two-Layer ReLU Networks
by: Arjevani, Yossi
Published: (2023)