Saved in:
| Main Authors: | Zhao, Jim, Singh, Sidak Pal, Lucchi, Aurelien |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2411.02139 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cubic regularized subspace Newton for non-convex optimization
by: Zhao, Jim, et al.
Published: (2024)
by: Zhao, Jim, et al.
Published: (2024)
Some Fundamental Aspects about Lipschitz Continuity of Neural Networks
by: Khromov, Grigory, et al.
Published: (2023)
by: Khromov, Grigory, et al.
Published: (2023)
Accelerating Neural Network Training Along Sharp and Flat Directions
by: Zakarin, Daniyar, et al.
Published: (2025)
by: Zakarin, Daniyar, et al.
Published: (2025)
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
by: Ormaniec, Weronika, et al.
Published: (2024)
by: Ormaniec, Weronika, et al.
Published: (2024)
Optimizer choice matters for the emergence of Neural Collapse
by: Zhao, Jim, et al.
Published: (2026)
by: Zhao, Jim, et al.
Published: (2026)
Hallmarks of Optimization Trajectories in Neural Networks: Directional Exploration and Redundancy
by: Singh, Sidak Pal, et al.
Published: (2024)
by: Singh, Sidak Pal, et al.
Published: (2024)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023)
by: Bozic, Vukasin, et al.
Published: (2023)
A Theoretical Analysis of the Learning Dynamics under Class Imbalance
by: Francazi, Emanuele, et al.
Published: (2022)
by: Francazi, Emanuele, et al.
Published: (2022)
Why Do We Need Warm-up? A Theoretical Perspective
by: Alimisis, Foivos, et al.
Published: (2025)
by: Alimisis, Foivos, et al.
Published: (2025)
Loss Landscape Characterization of Neural Networks without Over-Parametrization
by: Islamov, Rustem, et al.
Published: (2024)
by: Islamov, Rustem, et al.
Published: (2024)
Regularized Gauss-Newton for Optimizing Overparameterized Neural Networks
by: Adeoye, Adeyemi D., et al.
Published: (2024)
by: Adeoye, Adeyemi D., et al.
Published: (2024)
Exact Gauss-Newton Optimization for Training Deep Neural Networks
by: Korbit, Mikalai, et al.
Published: (2024)
by: Korbit, Mikalai, et al.
Published: (2024)
Avoiding spurious sharpness minimization broadens applicability of SAM
by: Singh, Sidak Pal, et al.
Published: (2025)
by: Singh, Sidak Pal, et al.
Published: (2025)
Local vs Global continual learning
by: Lanzillotta, Giulia, et al.
Published: (2024)
by: Lanzillotta, Giulia, et al.
Published: (2024)
Approximating Full Conformal Prediction for Neural Network Regression with Gauss-Newton Influence
by: Tailor, Dharmesh, et al.
Published: (2025)
by: Tailor, Dharmesh, et al.
Published: (2025)
Selective Forgetting in Option Calibration: An Operator-Theoretic Gauss-Newton Framework
by: Özsoy, Ahmet Umur
Published: (2025)
by: Özsoy, Ahmet Umur
Published: (2025)
Initial Guessing Bias: How Untrained Networks Favor Some Classes
by: Francazi, Emanuele, et al.
Published: (2023)
by: Francazi, Emanuele, et al.
Published: (2023)
Landscaping Linear Mode Connectivity
by: Singh, Sidak Pal, et al.
Published: (2024)
by: Singh, Sidak Pal, et al.
Published: (2024)
Generalized Linear Mode Connectivity for Transformers
by: Theus, Alexander, et al.
Published: (2025)
by: Theus, Alexander, et al.
Published: (2025)
Transformer Fusion with Optimal Transport
by: Imfeld, Moritz, et al.
Published: (2023)
by: Imfeld, Moritz, et al.
Published: (2023)
Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
by: Compagnoni, Enea Monzio, et al.
Published: (2024)
Gauss-Newton Unlearning for the LLM Era
by: McKinney, Lev, et al.
Published: (2026)
by: McKinney, Lev, et al.
Published: (2026)
Error whitening: Why Gauss-Newton outperforms Newton
by: McKay, Maricela Best, et al.
Published: (2026)
by: McKay, Maricela Best, et al.
Published: (2026)
A Riemannian Optimization Perspective of the Gauss-Newton Method for Feedforward Neural Networks
by: Cayci, Semih
Published: (2024)
by: Cayci, Semih
Published: (2024)
Fast Gauss-Newton for Multiclass Cross-Entropy
by: Korbit, Mikalai, et al.
Published: (2026)
by: Korbit, Mikalai, et al.
Published: (2026)
Towards Meta-Pruning via Optimal Transport
by: Theus, Alexander, et al.
Published: (2024)
by: Theus, Alexander, et al.
Published: (2024)
Characterizing Overfitting in Kernel Ridgeless Regression Through the Eigenspectrum
by: Cheng, Tin Sum, et al.
Published: (2024)
by: Cheng, Tin Sum, et al.
Published: (2024)
A Comprehensive Analysis on the Learning Curve in Kernel Ridge Regression
by: Cheng, Tin Sum, et al.
Published: (2024)
by: Cheng, Tin Sum, et al.
Published: (2024)
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
Incremental Gauss-Newton Descent for Machine Learning
by: Korbit, Mikalai, et al.
Published: (2024)
by: Korbit, Mikalai, et al.
Published: (2024)
A Structure-Guided Gauss-Newton Method for Shallow ReLU Neural Network
by: Cai, Zhiqiang, et al.
Published: (2024)
by: Cai, Zhiqiang, et al.
Published: (2024)
Unbiased and Sign Compression in Distributed Learning: Comparing Noise Resilience via SDEs
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
by: Compagnoni, Enea Monzio, et al.
Published: (2025)
Gradient Scalability and Taylor Surrogation of Quantum Cost Landscapes
by: Meyer, Sabri, et al.
Published: (2025)
by: Meyer, Sabri, et al.
Published: (2025)
Model Fusion via Retrofitting
by: Luenam, Phoomraphee, et al.
Published: (2025)
by: Luenam, Phoomraphee, et al.
Published: (2025)
Optimization Guarantees for Square-Root Natural-Gradient Variational Inference
by: Kumar, Navish, et al.
Published: (2025)
by: Kumar, Navish, et al.
Published: (2025)
Incremental Gauss--Newton Methods with Superlinear Convergence Rates
by: Zhou, Zhiling, et al.
Published: (2024)
by: Zhou, Zhiling, et al.
Published: (2024)
Gauss-Newton Natural Gradient Descent for Shape Learning
by: King, James, et al.
Published: (2026)
by: King, James, et al.
Published: (2026)
A Gauss-Newton Approach for Min-Max Optimization in Generative Adversarial Networks
by: Mishra, Neel, et al.
Published: (2024)
by: Mishra, Neel, et al.
Published: (2024)
Learning Morphisms with Gauss-Newton Approximation for Growing Networks
by: Lawton, Neal, et al.
Published: (2024)
by: Lawton, Neal, et al.
Published: (2024)
Double Momentum and Error Feedback for Clipping with Fast Rates and Differential Privacy
by: Islamov, Rustem, et al.
Published: (2025)
by: Islamov, Rustem, et al.
Published: (2025)
Similar Items
-
Cubic regularized subspace Newton for non-convex optimization
by: Zhao, Jim, et al.
Published: (2024) -
Some Fundamental Aspects about Lipschitz Continuity of Neural Networks
by: Khromov, Grigory, et al.
Published: (2023) -
Accelerating Neural Network Training Along Sharp and Flat Directions
by: Zakarin, Daniyar, et al.
Published: (2025) -
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
by: Ormaniec, Weronika, et al.
Published: (2024) -
Optimizer choice matters for the emergence of Neural Collapse
by: Zhao, Jim, et al.
Published: (2026)