On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning Regime
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Shuai, Voronin, Alexey, Cyr, Eric, Southworth, Ben |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Kourkoutas-Beta: A Sunspike-Driven Adam Optimizer with Desert Flair
von: Kassinos, Stavros C.
Veröffentlicht: (2025)
von: Kassinos, Stavros C.
Veröffentlicht: (2025)
Local properties of neural networks through the lens of layer-wise Hessians
von: Bolshim, Maxim, et al.
Veröffentlicht: (2025)
von: Bolshim, Maxim, et al.
Veröffentlicht: (2025)
ZetA: A Riemann Zeta-Scaled Extension of Adam for Deep Learning
von: BC, Samiksha
Veröffentlicht: (2025)
von: BC, Samiksha
Veröffentlicht: (2025)
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum
von: Zhang, Minxin, et al.
Veröffentlicht: (2026)
von: Zhang, Minxin, et al.
Veröffentlicht: (2026)
CAO: Curvature-Adaptive Optimization via Periodic Low-Rank Hessian Sketching
von: Du, Wenzhang
Veröffentlicht: (2025)
von: Du, Wenzhang
Veröffentlicht: (2025)
Explainable Attention-Based LSTM Framework for Early Detection of AI-Assisted Ransomware via File System Behavioral Analysis
von: Nayak, Prabhudarshi, et al.
Veröffentlicht: (2026)
von: Nayak, Prabhudarshi, et al.
Veröffentlicht: (2026)
Layer-Parallel Training for Transformers
von: Jiang, Shuai, et al.
Veröffentlicht: (2026)
von: Jiang, Shuai, et al.
Veröffentlicht: (2026)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
Benchmarking Generative AI Against Bayesian Optimization for Constrained Multi-Objective Inverse Design
von: Awan, Muhammad Bilal, et al.
Veröffentlicht: (2025)
von: Awan, Muhammad Bilal, et al.
Veröffentlicht: (2025)
Data-induced multiscale losses and efficient multirate gradient descent schemes
von: He, Juncai, et al.
Veröffentlicht: (2024)
von: He, Juncai, et al.
Veröffentlicht: (2024)
NeurOptimisation: The Spiking Way to Evolve
von: Cruz-Duarte, Jorge Mario, et al.
Veröffentlicht: (2025)
von: Cruz-Duarte, Jorge Mario, et al.
Veröffentlicht: (2025)
"Abuse Risks are Often Inherent to Product Features": Exploring AI Vendors' Bug Bounty and Responsible Disclosure Policies
von: Piao, Yangheran, et al.
Veröffentlicht: (2025)
von: Piao, Yangheran, et al.
Veröffentlicht: (2025)
i-DEQ: A stable inertial deep equilibrium model for image restoration
von: Clerc, Antonin, et al.
Veröffentlicht: (2026)
von: Clerc, Antonin, et al.
Veröffentlicht: (2026)
EB-gMCR: Energy-Based Generative Modeling for Signal Unmixing and Multivariate Curve Resolution
von: Chang, Yu-Tang, et al.
Veröffentlicht: (2025)
von: Chang, Yu-Tang, et al.
Veröffentlicht: (2025)
Refining Graphical Neural Network Predictions Using Flow Matching for Optimal Power Flow with Constraint-Satisfaction Guarantee
von: Khanal, Kshitiz
Veröffentlicht: (2025)
von: Khanal, Kshitiz
Veröffentlicht: (2025)
A Dual-Path Generative Framework for Zero-Day Fraud Detection in Banking Systems
von: Ismail, Nasim Abdirahman, et al.
Veröffentlicht: (2026)
von: Ismail, Nasim Abdirahman, et al.
Veröffentlicht: (2026)
Evaluation of Differential Privacy Mechanisms on Federated Learning
von: Varsani, Tejash
Veröffentlicht: (2025)
von: Varsani, Tejash
Veröffentlicht: (2025)
Sparsifying dimensionality reduction of PDE solution data with Bregman learning
von: Heeringa, Tjeerd Jan, et al.
Veröffentlicht: (2024)
von: Heeringa, Tjeerd Jan, et al.
Veröffentlicht: (2024)
Sparse Training of Neural Networks based on Multilevel Mirror Descent
von: Lunk, Yannick, et al.
Veröffentlicht: (2026)
von: Lunk, Yannick, et al.
Veröffentlicht: (2026)
FlowAdam: Implicit Regularization via Geometry-Aware Soft Momentum Injection
von: Singh, Devender, et al.
Veröffentlicht: (2026)
von: Singh, Devender, et al.
Veröffentlicht: (2026)
Physics Informed Differentiable Solvers for Learning Parametric Solution Manifolds in Heterogeneous Physical Systems
von: Panahi, Milad, et al.
Veröffentlicht: (2026)
von: Panahi, Milad, et al.
Veröffentlicht: (2026)
Deceptron: Learned Local Inverses for Fast and Stable Physics Inversion
von: Kachhadiya, Aaditya L.
Veröffentlicht: (2025)
von: Kachhadiya, Aaditya L.
Veröffentlicht: (2025)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
von: Hill, Brennen, et al.
Veröffentlicht: (2025)
von: Hill, Brennen, et al.
Veröffentlicht: (2025)
Deep Legendre Transform
von: Minabutdinov, Aleksey, et al.
Veröffentlicht: (2025)
von: Minabutdinov, Aleksey, et al.
Veröffentlicht: (2025)
Leveraging KANs for Expedient Training of Multichannel MLPs via Preconditioning and Geometric Refinement
von: Actor, Jonas A., et al.
Veröffentlicht: (2025)
von: Actor, Jonas A., et al.
Veröffentlicht: (2025)
Handbook of Convergence Theorems for (Stochastic) Gradient Methods
von: Garrigos, Guillaume, et al.
Veröffentlicht: (2023)
von: Garrigos, Guillaume, et al.
Veröffentlicht: (2023)
The Age of Sensorial Zero Trust: Why We Can No Longer Trust Our Senses
von: Xavier, Fabio Correa
Veröffentlicht: (2025)
von: Xavier, Fabio Correa
Veröffentlicht: (2025)
Gradient Descent as Implicit EM in Distance-Based Neural Models
von: Oursland, Alan
Veröffentlicht: (2025)
von: Oursland, Alan
Veröffentlicht: (2025)
Gradient descent provably escapes saddle points in the training of shallow ReLU networks
von: Cheridito, Patrick, et al.
Veröffentlicht: (2022)
von: Cheridito, Patrick, et al.
Veröffentlicht: (2022)
BreakFun: Jailbreaking LLMs via Schema Exploitation
von: Oskooei, Amirkia Rafiei, et al.
Veröffentlicht: (2025)
von: Oskooei, Amirkia Rafiei, et al.
Veröffentlicht: (2025)
Deep Learning and Elicitability for McKean-Vlasov FBSDEs With Common Noise
von: Antunes, Felipe J. P., et al.
Veröffentlicht: (2025)
von: Antunes, Felipe J. P., et al.
Veröffentlicht: (2025)
Total Generalized Variation regularization closes the gap between neural-eld and classical methods in seismic travel-time tomography
von: Kurosawa, Isao
Veröffentlicht: (2026)
von: Kurosawa, Isao
Veröffentlicht: (2026)
The Non-Linearity Perturbation Threshold: Width Scaling and Landscape Bifurcations in Deep Learning
von: Alexander, Michael
Veröffentlicht: (2026)
von: Alexander, Michael
Veröffentlicht: (2026)
Autoencoders in Function Space
von: Bunker, Justin, et al.
Veröffentlicht: (2024)
von: Bunker, Justin, et al.
Veröffentlicht: (2024)
Alternately-optimized SNN method for acoustic scattering problem in unbounded domain
von: Song, Haoming, et al.
Veröffentlicht: (2025)
von: Song, Haoming, et al.
Veröffentlicht: (2025)
Multirate Stein Variational Gradient Descent for Efficient Bayesian Sampling
von: Sarshar, Arash
Veröffentlicht: (2026)
von: Sarshar, Arash
Veröffentlicht: (2026)
Ghosts of Softmax: Complex Singularities That Limit Safe Step Sizes in Cross-Entropy
von: Sao, Piyush
Veröffentlicht: (2026)
von: Sao, Piyush
Veröffentlicht: (2026)
Context-dependent manifold learning: A neuromodulated constrained autoencoder approach
von: Adriaens, Jérôme, et al.
Veröffentlicht: (2026)
von: Adriaens, Jérôme, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Kourkoutas-Beta: A Sunspike-Driven Adam Optimizer with Desert Flair
von: Kassinos, Stavros C.
Veröffentlicht: (2025) -
Local properties of neural networks through the lens of layer-wise Hessians
von: Bolshim, Maxim, et al.
Veröffentlicht: (2025) -
ZetA: A Riemann Zeta-Scaled Extension of Adam for Deep Learning
von: BC, Samiksha
Veröffentlicht: (2025) -
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026) -
Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)