Series of Hessian-Vector Products for Tractable Saddle-Free Newton Optimisation of Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Oldewage, Elre T., Clarke, Ross M., Hernández-Lobato, José Miguel |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Studying K-FAC Heuristics by Viewing Adam through a Second-Order Lens
by: Clarke, Ross M., et al.
Published: (2023)
by: Clarke, Ross M., et al.
Published: (2023)
Warm Start Marginal Likelihood Optimisation for Iterative Gaussian Processes
by: Lin, Jihao Andreas, et al.
Published: (2024)
by: Lin, Jihao Andreas, et al.
Published: (2024)
Leveraging Task Structures for Improved Identifiability in Neural Network Representations
by: Chen, Wenlin, et al.
Published: (2023)
by: Chen, Wenlin, et al.
Published: (2023)
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes
by: Lin, Jihao Andreas, et al.
Published: (2024)
by: Lin, Jihao Andreas, et al.
Published: (2024)
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
by: Yamamoto, Naoya, et al.
Published: (2025)
by: Yamamoto, Naoya, et al.
Published: (2025)
Inertial Newton Algorithms Avoiding Strict Saddle Points
by: Castera, Camille
Published: (2021)
by: Castera, Camille
Published: (2021)
Training-Free Vector Quantization via Gaussian VAEs
by: Xu, Tongda, et al.
Published: (2025)
by: Xu, Tongda, et al.
Published: (2025)
Neural Network-based High-index Saddle Dynamics Method for Searching Saddle Points and Solution Landscape
by: Liu, Yuankai, et al.
Published: (2024)
by: Liu, Yuankai, et al.
Published: (2024)
Better Training Data Attribution via Better Inverse Hessian-Vector Products
by: Wang, Andrew, et al.
Published: (2025)
by: Wang, Andrew, et al.
Published: (2025)
Getting Free Bits Back from Rotational Symmetries in LLMs
by: He, Jiajun, et al.
Published: (2024)
by: He, Jiajun, et al.
Published: (2024)
Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks
by: Corlouer, Guillaume, et al.
Published: (2026)
by: Corlouer, Guillaume, et al.
Published: (2026)
Uncertainty Modeling in Graph Neural Networks via Stochastic Differential Equations
by: Bergna, Richard, et al.
Published: (2024)
by: Bergna, Richard, et al.
Published: (2024)
Dimension-Free Saddle-Point Escape in Muon
by: Long, Yanlin, et al.
Published: (2026)
by: Long, Yanlin, et al.
Published: (2026)
RECOMBINER: Robust and Enhanced Compression with Bayesian Implicit Neural Representations
by: He, Jiajun, et al.
Published: (2023)
by: He, Jiajun, et al.
Published: (2023)
Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle Escape
by: Bantzis, Ioannis, et al.
Published: (2025)
by: Bantzis, Ioannis, et al.
Published: (2025)
Towards Quantifying the Hessian Structure of Neural Networks
by: Dong, Zhaorui, et al.
Published: (2025)
by: Dong, Zhaorui, et al.
Published: (2025)
Online Newton Method for Bandit Convex Optimisation
by: Fokkema, Hidde, et al.
Published: (2024)
by: Fokkema, Hidde, et al.
Published: (2024)
Characterizing Learning in Deep Neural Networks using Tractable Algorithmic Complexity Analysis
by: Bakhtiarifard, Pedram, et al.
Published: (2026)
by: Bakhtiarifard, Pedram, et al.
Published: (2026)
Training Neural Samplers with Reverse Diffusive KL Divergence
by: He, Jiajun, et al.
Published: (2024)
by: He, Jiajun, et al.
Published: (2024)
Newton-CG methods for nonconvex unconstrained optimization with Hölder continuous Hessian
by: He, Chuan, et al.
Published: (2023)
by: He, Chuan, et al.
Published: (2023)
A Gaussian Process View on Observation Noise and Initialization in Wide Neural Networks
by: Calvo-Ordoñez, Sergio, et al.
Published: (2025)
by: Calvo-Ordoñez, Sergio, et al.
Published: (2025)
Exact, Tractable Gauss-Newton Optimization in Deep Reversible Architectures Reveal Poor Generalization
by: Buffelli, Davide, et al.
Published: (2024)
by: Buffelli, Davide, et al.
Published: (2024)
Causal Effect Estimation under Networked Interference without Networked Unconfoundedness Assumption
by: Chen, Weilin, et al.
Published: (2025)
by: Chen, Weilin, et al.
Published: (2025)
SGD with Partial Hessian for Deep Neural Networks Optimization
by: Sun, Ying, et al.
Published: (2024)
by: Sun, Ying, et al.
Published: (2024)
Diagnosing and fixing common problems in Bayesian optimization for molecule design
by: Tripp, Austin, et al.
Published: (2024)
by: Tripp, Austin, et al.
Published: (2024)
Stein Variational Newton Neural Network Ensembles
by: Flöge, Klemens, et al.
Published: (2024)
by: Flöge, Klemens, et al.
Published: (2024)
From Saddle Points Toward Global Minima: A Newton-Type Method on Wasserstein Space
by: Lascu, Razvan-Andrei, et al.
Published: (2026)
by: Lascu, Razvan-Andrei, et al.
Published: (2026)
Geometry of Critical Sets and Existence of Saddle Branches for Two-layer Neural Networks
by: Zhang, Leyang, et al.
Published: (2024)
by: Zhang, Leyang, et al.
Published: (2024)
Directional Convergence Near Small Initializations and Saddles in Two-Homogeneous Neural Networks
by: Kumar, Akshay, et al.
Published: (2024)
by: Kumar, Akshay, et al.
Published: (2024)
Loss Landscape of Shallow ReLU-like Neural Networks: Stationary Points, Saddle Escape, and Network Embedding
by: Wu, Frank Zhengqing, et al.
Published: (2024)
by: Wu, Frank Zhengqing, et al.
Published: (2024)
On Newton's Method to Unlearn Neural Networks
by: Bui, Nhung, et al.
Published: (2024)
by: Bui, Nhung, et al.
Published: (2024)
Tractable Probabilistic Graph Representation Learning with Graph-Induced Sum-Product Networks
by: Errica, Federico, et al.
Published: (2023)
by: Errica, Federico, et al.
Published: (2023)
Sum-Product-Set Networks: Deep Tractable Models for Tree-Structured Graphs
by: Papež, Milan, et al.
Published: (2024)
by: Papež, Milan, et al.
Published: (2024)
Decoupled PFNs: Identifiable Epistemic-Aleatoric Decomposition via Structured Synthetic Priors
by: Bergna, Richard, et al.
Published: (2026)
by: Bergna, Richard, et al.
Published: (2026)
There Was Never a Bottleneck in Concept Bottleneck Models
by: Almudévar, Antonio, et al.
Published: (2025)
by: Almudévar, Antonio, et al.
Published: (2025)
Post-Hoc Uncertainty Quantification in Pre-Trained Neural Networks via Activation-Level Gaussian Processes
by: Bergna, Richard, et al.
Published: (2025)
by: Bergna, Richard, et al.
Published: (2025)
Theoretical characterisation of the Gauss-Newton conditioning in Neural Networks
by: Zhao, Jim, et al.
Published: (2024)
by: Zhao, Jim, et al.
Published: (2024)
Hessian-Free Online Certified Unlearning
by: Qiao, Xinbao, et al.
Published: (2024)
by: Qiao, Xinbao, et al.
Published: (2024)
Revisit, Extend, and Enhance Hessian-Free Influence Functions
by: Yang, Ziao, et al.
Published: (2024)
by: Yang, Ziao, et al.
Published: (2024)
Similar Items
-
Studying K-FAC Heuristics by Viewing Adam through a Second-Order Lens
by: Clarke, Ross M., et al.
Published: (2023) -
Warm Start Marginal Likelihood Optimisation for Iterative Gaussian Processes
by: Lin, Jihao Andreas, et al.
Published: (2024) -
Leveraging Task Structures for Improved Identifiability in Neural Network Representations
by: Chen, Wenlin, et al.
Published: (2023) -
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
by: Zhang, Yedi, et al.
Published: (2025) -
Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes
by: Lin, Jihao Andreas, et al.
Published: (2024)