Loss Landscape Characterization of Neural Networks without Over-Parametrization
Fuente:
arXiv
Saved in:
| Main Authors: | Islamov, Rustem, Ajroldi, Niccolò, Orvieto, Antonio, Lucchi, Aurelien |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
Why Do We Need Warm-up? A Theoretical Perspective
by: Alimisis, Foivos, et al.
Published: (2025)
by: Alimisis, Foivos, et al.
Published: (2025)
Double Momentum and Error Feedback for Clipping with Fast Rates and Differential Privacy
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Safe-EF: Error Feedback for Nonsmooth Constrained Optimization
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
SDEs for Minimax Optimization
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Towards Faster Decentralized Stochastic Optimization with Communication Compression
by: Islamov, Rustem, et al.
Published: (2024)
by: Islamov, Rustem, et al.
Published: (2024)
An Adaptive Stochastic Gradient Method with Non-negative Gauss-Newton Stepsizes
by: Orvieto, Antonio, et al.
Published: (2024)
by: Orvieto, Antonio, et al.
Published: (2024)
Cubic regularized subspace Newton for non-convex optimization
by: Zhao, Jim, et al.
Published: (2024)
by: Zhao, Jim, et al.
Published: (2024)
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling
by: Srećković, Teodora, et al.
Published: (2025)
by: Srećković, Teodora, et al.
Published: (2025)
Recurrent neural networks: vanishing and exploding gradients are not the end of the story
by: Zucchet, Nicolas, et al.
Published: (2024)
by: Zucchet, Nicolas, et al.
Published: (2024)
Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
When, Where and Why to Average Weights?
by: Ajroldi, Niccolò, et al.
Published: (2025)
by: Ajroldi, Niccolò, et al.
Published: (2025)
Optimality-Informed Neural Networks for Solving Parametric Optimization Problems
by: Hoffmann, Matthias K., et al.
Published: (2025)
by: Hoffmann, Matthias K., et al.
Published: (2025)
Gradient Descent on Logistic Regression: Do Large Step-Sizes Work with Data on the Sphere?
by: Meng, Si Yi, et al.
Published: (2025)
by: Meng, Si Yi, et al.
Published: (2025)
Challenges in Training PINNs: A Loss Landscape Perspective
by: Rathore, Pratik, et al.
Published: (2024)
by: Rathore, Pratik, et al.
Published: (2024)
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
by: Meng, Si Yi, et al.
Published: (2024)
by: Meng, Si Yi, et al.
Published: (2024)
Optimization Over Trained Neural Networks: Taking a Relaxing Walk
by: Tong, Jiatai, et al.
Published: (2024)
by: Tong, Jiatai, et al.
Published: (2024)
Discrete-time Contraction-based Control of Nonlinear Systems with Parametric Uncertainties using Neural Networks
by: Wei, Lai, et al.
Published: (2021)
by: Wei, Lai, et al.
Published: (2021)
Parameter-Adaptive Approximate MPC: Tuning Neural-Network Controllers without Retraining
by: Hose, Henrik, et al.
Published: (2024)
by: Hose, Henrik, et al.
Published: (2024)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
by: Tanguy, Eloi
Published: (2023)
by: Tanguy, Eloi
Published: (2023)
Over-parameterised Shallow Neural Networks with Asymmetrical Node Scaling: Global Convergence Guarantees and Feature Learning
by: Caron, Francois, et al.
Published: (2023)
by: Caron, Francois, et al.
Published: (2023)
Regret-Optimal Federated Transfer Learning for Kernel Regression with Applications in American Option Pricing
by: Yang, Xuwei, et al.
Published: (2023)
by: Yang, Xuwei, et al.
Published: (2023)
Unbiased and Sign Compression in Distributed Learning: Comparing Noise Resilience via SDEs
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
Physics-Informed Neural Network Lyapunov Functions: PDE Characterization, Learning, and Verification
by: Liu, Jun, et al.
Published: (2023)
by: Liu, Jun, et al.
Published: (2023)
A Precise Characterization of SGD Stability Using Loss Surface Geometry
by: Dexter, Gregory, et al.
Published: (2024)
by: Dexter, Gregory, et al.
Published: (2024)
Geometric Foundations of Tuning without Forgetting in Neural ODEs
by: Bayram, Erkan, et al.
Published: (2025)
by: Bayram, Erkan, et al.
Published: (2025)
Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
by: Elhassan, Fay, et al.
Published: (2025)
by: Elhassan, Fay, et al.
Published: (2025)
The Optimiser Hidden in Plain Sight: Training with the Loss Landscape's Induced Metric
by: Harvey, Thomas R.
Published: (2025)
by: Harvey, Thomas R.
Published: (2025)
Global Convergence of SGD For Logistic Loss on Two Layer Neural Nets
by: Gopalani, Pulkit, et al.
Published: (2023)
by: Gopalani, Pulkit, et al.
Published: (2023)
Nonlinear Dynamics In Optimization Landscape of Shallow Neural Networks with Tunable Leaky ReLU
by: Liu, Jingzhou
Published: (2025)
by: Liu, Jingzhou
Published: (2025)
Optimization Landscapes Learned: Proxy Networks Boost Convergence in Physics-based Inverse Problems
by: Goyal, Girnar, et al.
Published: (2025)
by: Goyal, Girnar, et al.
Published: (2025)
Parametrizing Convex Sets Using Sublinear Neural Networks
by: Martinet, Eloi
Published: (2026)
by: Martinet, Eloi
Published: (2026)
Parametric Nonconvex Optimization via Convex Surrogates
by: Wang, Renzi, et al.
Published: (2026)
by: Wang, Renzi, et al.
Published: (2026)
KKT-Informed Neural Network
by: Femine, Carmine Delle
Published: (2024)
by: Femine, Carmine Delle
Published: (2024)
Optimal Depth of Neural Networks
by: Qi, Qian
Published: (2025)
by: Qi, Qian
Published: (2025)
Learning Algorithm Hyperparameters for Fast Parametric Convex Optimization
by: Sambharya, Rajiv, et al.
Published: (2024)
by: Sambharya, Rajiv, et al.
Published: (2024)
Similar Items
-
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
by: Islamov, Rustem, et al.
Published: (2025) -
Why Do We Need Warm-up? A Theoretical Perspective
by: Alimisis, Foivos, et al.
Published: (2025) -
Double Momentum and Error Feedback for Clipping with Fast Rates and Differential Privacy
by: Islamov, Rustem, et al.
Published: (2025) -
On the Role of Batch Size in Stochastic Conditional Gradient Methods
by: Islamov, Rustem, et al.
Published: (2026) -
Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions
by: Islamov, Rustem, et al.
Published: (2026)