How Does the ReLU Activation Affect the Implicit Bias of Gradient Descent on High-dimensional Neural Network Regression?
Fuente:
arXiv
Saved in:
| Main Authors: | Lai, Kuo-Wei, Wang, Guanghui, Tao, Molei, Muthukumar, Vidya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
General Loss Functions Lead to (Approximate) Interpolation in High Dimensions
by: Lai, Kuo-Wei, et al.
Published: (2023)
by: Lai, Kuo-Wei, et al.
Published: (2023)
Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
by: Min, Hancheng, et al.
Published: (2025)
by: Min, Hancheng, et al.
Published: (2025)
Convex Formulations for Training Two-Layer ReLU Neural Networks
by: Prakhya, Karthik, et al.
Published: (2024)
by: Prakhya, Karthik, et al.
Published: (2024)
Function Gradient Approximation with Random Shallow ReLU Networks with Control Applications
by: Lamperski, Andrew, et al.
Published: (2024)
by: Lamperski, Andrew, et al.
Published: (2024)
Stability and Performance Analysis of Discrete-Time ReLU Recurrent Neural Networks
by: Noori, Sahel Vahedi, et al.
Published: (2024)
by: Noori, Sahel Vahedi, et al.
Published: (2024)
Global Convergence of Policy Gradient Methods for ReLU Controllers in Linear Quadratic Regulation
by: Rodriguez-Gil, Jhojan A., et al.
Published: (2026)
by: Rodriguez-Gil, Jhojan A., et al.
Published: (2026)
Hidden Minima in Two-Layer ReLU Networks
by: Arjevani, Yossi
Published: (2023)
by: Arjevani, Yossi
Published: (2023)
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
by: Liao, Fangshuo, et al.
Published: (2023)
by: Liao, Fangshuo, et al.
Published: (2023)
Adversarial Training of Two-Layer Polynomial and ReLU Activation Networks via Convex Optimization
by: Kuelbs, Daniel, et al.
Published: (2024)
by: Kuelbs, Daniel, et al.
Published: (2024)
Convex Relaxations of ReLU Neural Networks Approximate Global Optima in Polynomial Time
by: Kim, Sungyoon, et al.
Published: (2024)
by: Kim, Sungyoon, et al.
Published: (2024)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
by: Cai, Yuhang, et al.
Published: (2025)
by: Cai, Yuhang, et al.
Published: (2025)
ReLU Surrogates in Mixed-Integer MPC for Irrigation Scheduling
by: Agyeman, Bernard T., et al.
Published: (2024)
by: Agyeman, Bernard T., et al.
Published: (2024)
ReLU Networks for Model Predictive Control: Network Complexity and Performance Guarantees
by: Li, Xingchen, et al.
Published: (2026)
by: Li, Xingchen, et al.
Published: (2026)
On the ReLU Lagrangian Cuts for Stochastic Mixed Integer Programming
by: Deng, Haoyun, et al.
Published: (2024)
by: Deng, Haoyun, et al.
Published: (2024)
LMI hierarchies for stability analysis of ReLU feedback systems
by: Magron, Victor, et al.
Published: (2024)
by: Magron, Victor, et al.
Published: (2024)
Homotopy Relaxation Training Algorithms for Infinite-Width Two-Layer ReLU Neural Networks
by: Yang, Yahong, et al.
Published: (2023)
by: Yang, Yahong, et al.
Published: (2023)
Computational Tradeoffs of Optimization-Based Bound Tightening in ReLU Networks
by: Badilla, Fabian, et al.
Published: (2023)
by: Badilla, Fabian, et al.
Published: (2023)
Nonlinear Dynamics In Optimization Landscape of Shallow Neural Networks with Tunable Leaky ReLU
by: Liu, Jingzhou
Published: (2025)
by: Liu, Jingzhou
Published: (2025)
Normalization of ReLU Dual for Cut Generation in Stochastic Mixed-Integer Programs
by: Bansal, Akul, et al.
Published: (2026)
by: Bansal, Akul, et al.
Published: (2026)
Approximation with Random Shallow ReLU Networks with Applications to Model Reference Adaptive Control
by: Lamperski, Andrew, et al.
Published: (2024)
by: Lamperski, Andrew, et al.
Published: (2024)
Geometry-induced Regularization in Deep ReLU Neural Networks
by: Bona-Pellissier, Joachim, et al.
Published: (2024)
by: Bona-Pellissier, Joachim, et al.
Published: (2024)
Why Smooth Stability Assumptions Fail for ReLU Learning
by: Katende, Ronald
Published: (2025)
by: Katende, Ronald
Published: (2025)
An analysis of optimization problems involving ReLU neural networks
by: Plate, Christoph, et al.
Published: (2025)
by: Plate, Christoph, et al.
Published: (2025)
Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification
by: Jung, Hyunji, et al.
Published: (2025)
by: Jung, Hyunji, et al.
Published: (2025)
A Complete Set of Quadratic Constraints for Repeated ReLU and Generalizations
by: Noori, Sahel Vahedi, et al.
Published: (2024)
by: Noori, Sahel Vahedi, et al.
Published: (2024)
Pruning for efficient deterministic global optimization over trained ReLU neural networks
by: Lastrucci, Giacomo, et al.
Published: (2026)
by: Lastrucci, Giacomo, et al.
Published: (2026)
An Efficient Alternating Algorithm for ReLU-based Symmetric Matrix Decomposition
by: Wang, Qingsong
Published: (2025)
by: Wang, Qingsong
Published: (2025)
Discrete-Time Stability Analysis of ReLU Feedback Systems via Integral Quadratic Constraints
by: Noori, Sahel Vahedi, et al.
Published: (2025)
by: Noori, Sahel Vahedi, et al.
Published: (2025)
MIQCQP reformulation of the ReLU neural networks Lipschitz constant estimation problem
by: Sbihi, Mohammed, et al.
Published: (2024)
by: Sbihi, Mohammed, et al.
Published: (2024)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
by: Xu, Xianliang, et al.
Published: (2024)
by: Xu, Xianliang, et al.
Published: (2024)
Provable Acceleration of Nesterov's Accelerated Gradient for Rectangular Matrix Factorization and Linear Neural Networks
by: Xu, Zhenghao, et al.
Published: (2024)
by: Xu, Zhenghao, et al.
Published: (2024)
Path-conditioned training: a principled way to rescale ReLU neural networks
by: Lebeurrier, Arthur, et al.
Published: (2026)
by: Lebeurrier, Arthur, et al.
Published: (2026)
Non-Singularity of the Gradient Descent map for Neural Networks with Piecewise Analytic Activations
by: Crăciun, Alexandru, et al.
Published: (2025)
by: Crăciun, Alexandru, et al.
Published: (2025)
Local Lipschitz Constant Computation of ReLU-FNNs: Upper Bound Computation with Exactness Verification
by: Ebihara, Yoshio, et al.
Published: (2023)
by: Ebihara, Yoshio, et al.
Published: (2023)
Implicit Bias and Convergence of Matrix Stochastic Mirror Descent
by: Akhtiamov, Danil, et al.
Published: (2026)
by: Akhtiamov, Danil, et al.
Published: (2026)
Implicit Bias of Spectral Descent and Muon on Multiclass Separable Data
by: Fan, Chen, et al.
Published: (2025)
by: Fan, Chen, et al.
Published: (2025)
Understanding the Implicit Regularization of Gradient Descent in Over-parameterized Models
by: Ma, Jianhao, et al.
Published: (2025)
by: Ma, Jianhao, et al.
Published: (2025)
High-order Accumulative Regularization for Gradient Minimization in Convex Programming
by: Ji, Yao, et al.
Published: (2025)
by: Ji, Yao, et al.
Published: (2025)
Constructive Universal Approximation and Finite Sample Memorization by Narrow Deep ReLU Networks
by: Hernández, Martín, et al.
Published: (2024)
by: Hernández, Martín, et al.
Published: (2024)
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
by: Sato, Naoki, et al.
Published: (2024)
by: Sato, Naoki, et al.
Published: (2024)
Similar Items
-
General Loss Functions Lead to (Approximate) Interpolation in High Dimensions
by: Lai, Kuo-Wei, et al.
Published: (2023) -
Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
by: Min, Hancheng, et al.
Published: (2025) -
Convex Formulations for Training Two-Layer ReLU Neural Networks
by: Prakhya, Karthik, et al.
Published: (2024) -
Function Gradient Approximation with Random Shallow ReLU Networks with Control Applications
by: Lamperski, Andrew, et al.
Published: (2024) -
Stability and Performance Analysis of Discrete-Time ReLU Recurrent Neural Networks
by: Noori, Sahel Vahedi, et al.
Published: (2024)