Towards Quantifying the Hessian Structure of Neural Networks
Fuente:
arXiv
Salvato in:
| Autori principali: | Dong, Zhaorui, Zhang, Yushun, Yao, Jianfeng, Sun, Ruoyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SGD with Partial Hessian for Deep Neural Networks Optimization
di: Sun, Ying, et al.
Pubblicazione: (2024)
di: Sun, Ying, et al.
Pubblicazione: (2024)
Adam Converges Without Any Modification On Update Rules
di: Zhang, Yushun, et al.
Pubblicazione: (2026)
di: Zhang, Yushun, et al.
Pubblicazione: (2026)
Provable Adaptivity of Adam under Non-uniform Smoothness
di: Wang, Bohan, et al.
Pubblicazione: (2022)
di: Wang, Bohan, et al.
Pubblicazione: (2022)
Constrained Bi-Level Optimization: Proximal Lagrangian Value function Approach and Hessian-free Algorithm
di: Yao, Wei, et al.
Pubblicazione: (2024)
di: Yao, Wei, et al.
Pubblicazione: (2024)
Moreau Envelope for Nonconvex Bi-Level Optimization: A Single-loop and Hessian-free Solution Strategy
di: Liu, Risheng, et al.
Pubblicazione: (2024)
di: Liu, Risheng, et al.
Pubblicazione: (2024)
Continuous Q-Score Matching: Diffusion Guided Reinforcement Learning for Continuous-Time Control
di: Hua, Chengxiu, et al.
Pubblicazione: (2025)
di: Hua, Chengxiu, et al.
Pubblicazione: (2025)
Stochastic Hessian Fittings with Lie Groups
di: Li, Xi-Lin
Pubblicazione: (2024)
di: Li, Xi-Lin
Pubblicazione: (2024)
A Hessian-Aware Stochastic Differential Equation for Modelling SGD
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
di: Semenov, Andrei, et al.
Pubblicazione: (2025)
di: Semenov, Andrei, et al.
Pubblicazione: (2025)
A Single-Loop Smoothed Gradient Descent-Ascent Algorithm for Nonconvex-Concave Min-Max Problems
di: Zhang, Jiawei, et al.
Pubblicazione: (2020)
di: Zhang, Jiawei, et al.
Pubblicazione: (2020)
Optimal Hessian/Jacobian-Free Nonconvex-PL Bilevel Optimization
di: Huang, Feihu
Pubblicazione: (2024)
di: Huang, Feihu
Pubblicazione: (2024)
State estimations and noise identifications with intermittent corrupted observations via Bayesian variational inference
di: Sun, Peng, et al.
Pubblicazione: (2026)
di: Sun, Peng, et al.
Pubblicazione: (2026)
Towards Guided Descent: Optimization Algorithms for Training Neural Networks At Scale
di: Nagwekar, Ansh
Pubblicazione: (2025)
di: Nagwekar, Ansh
Pubblicazione: (2025)
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
di: Yamamoto, Naoya, et al.
Pubblicazione: (2025)
di: Yamamoto, Naoya, et al.
Pubblicazione: (2025)
Newton-CG methods for nonconvex unconstrained optimization with Hölder continuous Hessian
di: He, Chuan, et al.
Pubblicazione: (2023)
di: He, Chuan, et al.
Pubblicazione: (2023)
Towards Understanding Gradient Flow Dynamics of Homogeneous Neural Networks Beyond the Origin
di: Kumar, Akshay, et al.
Pubblicazione: (2025)
di: Kumar, Akshay, et al.
Pubblicazione: (2025)
Towards Optimal Branching of Linear and Semidefinite Relaxations for Neural Network Robustness Certification
di: Anderson, Brendon G., et al.
Pubblicazione: (2021)
di: Anderson, Brendon G., et al.
Pubblicazione: (2021)
Global Convergence of Natural Policy Gradient with Hessian-aided Momentum Variance Reduction
di: Feng, Jie, et al.
Pubblicazione: (2024)
di: Feng, Jie, et al.
Pubblicazione: (2024)
On the Convergence of Adam under Non-uniform Smoothness: Separability from SGDM and Beyond
di: Wang, Bohan, et al.
Pubblicazione: (2024)
di: Wang, Bohan, et al.
Pubblicazione: (2024)
The Jacobian and Hessian of the Kullback-Leibler Divergence between Multivariate Gaussian Distributions (Technical Report)
di: Maroñas, Juan
Pubblicazione: (2025)
di: Maroñas, Juan
Pubblicazione: (2025)
Online estimation of the inverse of the Hessian for stochastic optimization with application to universal stochastic Newton algorithms
di: Godichon-Baggioni, Antoine, et al.
Pubblicazione: (2024)
di: Godichon-Baggioni, Antoine, et al.
Pubblicazione: (2024)
Towards Quantifying the Preconditioning Effect of Adam
di: Das, Rudrajit, et al.
Pubblicazione: (2024)
di: Das, Rudrajit, et al.
Pubblicazione: (2024)
Convergence of stochastic gradient descent under a local Lojasiewicz condition for deep neural networks
di: An, Jing, et al.
Pubblicazione: (2023)
di: An, Jing, et al.
Pubblicazione: (2023)
Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity Limits
di: Sadiev, Abdurakhmon, et al.
Pubblicazione: (2025)
di: Sadiev, Abdurakhmon, et al.
Pubblicazione: (2025)
Achieving ${O}(ε^{-1.5})$ Complexity in Hessian/Jacobian-free Stochastic Bilevel Optimization
di: Yang, Yifan, et al.
Pubblicazione: (2023)
di: Yang, Yifan, et al.
Pubblicazione: (2023)
(Almost) Smooth Sailing: Towards Numerical Stability of Neural Networks Through Differentiable Regularization of the Condition Number
di: Nenov, Rossen, et al.
Pubblicazione: (2024)
di: Nenov, Rossen, et al.
Pubblicazione: (2024)
Accelerated Stochastic ExtraGradient: Mixing Hessian and Gradient Similarity to Reduce Communication in Distributed and Federated Learning
di: Bylinkin, Dmitry, et al.
Pubblicazione: (2024)
di: Bylinkin, Dmitry, et al.
Pubblicazione: (2024)
Feature Augmentation of GNNs for ILPs: Local Uniqueness Suffices
di: Han, Qingyu, et al.
Pubblicazione: (2025)
di: Han, Qingyu, et al.
Pubblicazione: (2025)
Regularized Gradient Clipping Provably Trains Wide and Deep Neural Networks
di: Tucat, Matteo, et al.
Pubblicazione: (2024)
di: Tucat, Matteo, et al.
Pubblicazione: (2024)
NeST-BO: Fast Local Bayesian Optimization via Newton-Step Targeting of Gradient and Hessian Information
di: Tang, Wei-Ting, et al.
Pubblicazione: (2025)
di: Tang, Wei-Ting, et al.
Pubblicazione: (2025)
Regularized Adaptive Momentum Dual Averaging with an Efficient Inexact Subproblem Solver for Training Structured Neural Network
di: Huang, Zih-Syuan, et al.
Pubblicazione: (2024)
di: Huang, Zih-Syuan, et al.
Pubblicazione: (2024)
A Hessian-Free Actor-Critic Algorithm for Bi-Level Reinforcement Learning with Applications to LLM Fine-Tuning
di: Zeng, Sihan, et al.
Pubblicazione: (2026)
di: Zeng, Sihan, et al.
Pubblicazione: (2026)
Quadratic Gradient: A Unified Framework Bridging Gradient Descent and Newton-Type Methods by Synthesizing Hessians and Gradients
di: Chiang, John
Pubblicazione: (2022)
di: Chiang, John
Pubblicazione: (2022)
First-ish Order Methods: Hessian-aware Scalings of Gradient Descent
di: Smee, Oscar, et al.
Pubblicazione: (2025)
di: Smee, Oscar, et al.
Pubblicazione: (2025)
Geometry of Critical Sets and Existence of Saddle Branches for Two-layer Neural Networks
di: Zhang, Leyang, et al.
Pubblicazione: (2024)
di: Zhang, Leyang, et al.
Pubblicazione: (2024)
Active Learning of Deep Neural Networks via Gradient-Free Cutting Planes
di: Zhang, Erica, et al.
Pubblicazione: (2024)
di: Zhang, Erica, et al.
Pubblicazione: (2024)
Optimal Depth of Neural Networks
di: Qi, Qian
Pubblicazione: (2025)
di: Qi, Qian
Pubblicazione: (2025)
KKT-Informed Neural Network
di: Femine, Carmine Delle
Pubblicazione: (2024)
di: Femine, Carmine Delle
Pubblicazione: (2024)
Efficient Reachability Analysis for Convolutional Neural Networks Using Hybrid Zonotopes
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
Analyzing Neural Network-Based Generative Diffusion Models through Convex Optimization
di: Zhang, Fangzhao, et al.
Pubblicazione: (2024)
di: Zhang, Fangzhao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SGD with Partial Hessian for Deep Neural Networks Optimization
di: Sun, Ying, et al.
Pubblicazione: (2024) -
Adam Converges Without Any Modification On Update Rules
di: Zhang, Yushun, et al.
Pubblicazione: (2026) -
Provable Adaptivity of Adam under Non-uniform Smoothness
di: Wang, Bohan, et al.
Pubblicazione: (2022) -
Constrained Bi-Level Optimization: Proximal Lagrangian Value function Approach and Hessian-free Algorithm
di: Yao, Wei, et al.
Pubblicazione: (2024) -
Moreau Envelope for Nonconvex Bi-Level Optimization: A Single-loop and Hessian-free Solution Strategy
di: Liu, Risheng, et al.
Pubblicazione: (2024)