Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
Fuente:
arXiv
Saved in:
| Main Authors: | Bolshim, Maxim, Kugaevskikh, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Local properties of neural networks through the lens of layer-wise Hessians
by: Bolshim, Maxim, et al.
Published: (2025)
by: Bolshim, Maxim, et al.
Published: (2025)
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
by: Bolshim, Maxim, et al.
Published: (2026)
by: Bolshim, Maxim, et al.
Published: (2026)
CAO: Curvature-Adaptive Optimization via Periodic Low-Rank Hessian Sketching
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
ZetA: A Riemann Zeta-Scaled Extension of Adam for Deep Learning
by: BC, Samiksha
Published: (2025)
by: BC, Samiksha
Published: (2025)
Refining Graphical Neural Network Predictions Using Flow Matching for Optimal Power Flow with Constraint-Satisfaction Guarantee
by: Khanal, Kshitiz
Published: (2025)
by: Khanal, Kshitiz
Published: (2025)
NOVAK: Unified adaptive optimizer for deep neural networks
by: Kavun, Sergii
Published: (2026)
by: Kavun, Sergii
Published: (2026)
On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning Regime
by: Jiang, Shuai, et al.
Published: (2026)
by: Jiang, Shuai, et al.
Published: (2026)
NeurOptimisation: The Spiking Way to Evolve
by: Cruz-Duarte, Jorge Mario, et al.
Published: (2025)
by: Cruz-Duarte, Jorge Mario, et al.
Published: (2025)
Deep Legendre Transform
by: Minabutdinov, Aleksey, et al.
Published: (2025)
by: Minabutdinov, Aleksey, et al.
Published: (2025)
An in-depth look at approximation via deep and narrow neural networks
by: Dommel, Joris, et al.
Published: (2025)
by: Dommel, Joris, et al.
Published: (2025)
Predicting Traffic Accident Severity with Deep Neural Networks
by: Bibb, Meghan, et al.
Published: (2025)
by: Bibb, Meghan, et al.
Published: (2025)
TED++: Submanifold-Aware Backdoor Detection via Layerwise Tubular-Neighbourhood Screening
by: Le, Nam, et al.
Published: (2025)
by: Le, Nam, et al.
Published: (2025)
Kourkoutas-Beta: A Sunspike-Driven Adam Optimizer with Desert Flair
by: Kassinos, Stavros C.
Published: (2025)
by: Kassinos, Stavros C.
Published: (2025)
Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning
by: Troxell, David, et al.
Published: (2026)
by: Troxell, David, et al.
Published: (2026)
Explicit Dropout: Deterministic Regularization for Transformer Architectures
by: Agrawal, Vidhi, et al.
Published: (2026)
by: Agrawal, Vidhi, et al.
Published: (2026)
Massive Redundancy in Gradient Transport Enables Sparse Online Learning
by: Merin, Aur Shalev
Published: (2026)
by: Merin, Aur Shalev
Published: (2026)
Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
by: Ho, Siu Hang, et al.
Published: (2025)
by: Ho, Siu Hang, et al.
Published: (2025)
Is Cambodia the World's Largest Cashew Producer?
by: Chaya, Veasna, et al.
Published: (2024)
by: Chaya, Veasna, et al.
Published: (2024)
i-DEQ: A stable inertial deep equilibrium model for image restoration
by: Clerc, Antonin, et al.
Published: (2026)
by: Clerc, Antonin, et al.
Published: (2026)
Functional Similarity Metric for Neural Networks: Overcoming Parametric Ambiguity via Activation Region Analysis
by: Hennadii, Kutomanov
Published: (2026)
by: Hennadii, Kutomanov
Published: (2026)
Escaping Saddle Points via Curvature-Calibrated Perturbations: A Complete Analysis with Explicit Constants and Empirical Validation
by: Alpay, Faruk, et al.
Published: (2025)
by: Alpay, Faruk, et al.
Published: (2025)
VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks
by: Johnson, David R., et al.
Published: (2025)
by: Johnson, David R., et al.
Published: (2025)
H-Model: Dynamic Neural Architectures for Adaptive Processing
by: Hospodarchuk, Dmytro
Published: (2025)
by: Hospodarchuk, Dmytro
Published: (2025)
A Hybrid Inductive-Transductive Network for Traffic Flow Imputation on Unsampled Locations
by: Rahimiasl, Mohammadmahdi, et al.
Published: (2025)
by: Rahimiasl, Mohammadmahdi, et al.
Published: (2025)
The Normalized Difference Layer: A Differentiable Spectral Index Formulation for Deep Learning
by: Lotfi, Ali, et al.
Published: (2026)
by: Lotfi, Ali, et al.
Published: (2026)
CellARC: Measuring Intelligence with Cellular Automata
by: Lžičař, Miroslav
Published: (2025)
by: Lžičař, Miroslav
Published: (2025)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
by: Dang, Kieu, et al.
Published: (2025)
by: Dang, Kieu, et al.
Published: (2025)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
by: Berthier, Louis, et al.
Published: (2025)
by: Berthier, Louis, et al.
Published: (2025)
Tricks and Plug-ins for Gradient Boosting with Transformers
by: Fang, Biyi, et al.
Published: (2025)
by: Fang, Biyi, et al.
Published: (2025)
Benchmarking Generative AI Against Bayesian Optimization for Constrained Multi-Objective Inverse Design
by: Awan, Muhammad Bilal, et al.
Published: (2025)
by: Awan, Muhammad Bilal, et al.
Published: (2025)
Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum
by: Zhang, Minxin, et al.
Published: (2026)
by: Zhang, Minxin, et al.
Published: (2026)
CVCM Track Circuits Pre-emptive Failure Diagnostics for Predictive Maintenance Using Deep Neural Networks
by: Mukherjee, Debdeep, et al.
Published: (2025)
by: Mukherjee, Debdeep, et al.
Published: (2025)
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
by: Loza, Andrew J., et al.
Published: (2025)
by: Loza, Andrew J., et al.
Published: (2025)
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
by: Mysore, Naveen
Published: (2026)
by: Mysore, Naveen
Published: (2026)
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
by: Alnemari, Mohammed, et al.
Published: (2026)
by: Alnemari, Mohammed, et al.
Published: (2026)
Machine Unlearning for Class Removal through SISA-based Deep Neural Network Architectures
by: Mahi, Ishrak Hamim, et al.
Published: (2026)
by: Mahi, Ishrak Hamim, et al.
Published: (2026)
A Dual-Path Generative Framework for Zero-Day Fraud Detection in Banking Systems
by: Ismail, Nasim Abdirahman, et al.
Published: (2026)
by: Ismail, Nasim Abdirahman, et al.
Published: (2026)
Sparse Training of Neural Networks based on Multilevel Mirror Descent
by: Lunk, Yannick, et al.
Published: (2026)
by: Lunk, Yannick, et al.
Published: (2026)
NetworkNet: A Deep Neural Network Approach for Random Networks with Sparse Nodal Attributes and Complex Nodal Heterogeneity
by: Xing, Zhaoyu, et al.
Published: (2026)
by: Xing, Zhaoyu, et al.
Published: (2026)
Complex-valued convolutional neural network classification of hand gesture from radar images
by: Khandan, Shokooh
Published: (2024)
by: Khandan, Shokooh
Published: (2024)
Similar Items
-
Local properties of neural networks through the lens of layer-wise Hessians
by: Bolshim, Maxim, et al.
Published: (2025) -
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
by: Bolshim, Maxim, et al.
Published: (2026) -
CAO: Curvature-Adaptive Optimization via Periodic Low-Rank Hessian Sketching
by: Du, Wenzhang
Published: (2025) -
ZetA: A Riemann Zeta-Scaled Extension of Adam for Deep Learning
by: BC, Samiksha
Published: (2025) -
Refining Graphical Neural Network Predictions Using Flow Matching for Optimal Power Flow with Constraint-Satisfaction Guarantee
by: Khanal, Kshitiz
Published: (2025)