Series of Hessian-Vector Products for Tractable Saddle-Free Newton Optimisation of Neural Networks
Fuente:
arXiv
Guardado en:
| Autores principales: | Oldewage, Elre T., Clarke, Ross M., Hernández-Lobato, José Miguel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Studying K-FAC Heuristics by Viewing Adam through a Second-Order Lens
por: Clarke, Ross M., et al.
Publicado: (2023)
por: Clarke, Ross M., et al.
Publicado: (2023)
Warm Start Marginal Likelihood Optimisation for Iterative Gaussian Processes
por: Lin, Jihao Andreas, et al.
Publicado: (2024)
por: Lin, Jihao Andreas, et al.
Publicado: (2024)
Leveraging Task Structures for Improved Identifiability in Neural Network Representations
por: Chen, Wenlin, et al.
Publicado: (2023)
por: Chen, Wenlin, et al.
Publicado: (2023)
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
por: Zhang, Yedi, et al.
Publicado: (2025)
por: Zhang, Yedi, et al.
Publicado: (2025)
Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes
por: Lin, Jihao Andreas, et al.
Publicado: (2024)
por: Lin, Jihao Andreas, et al.
Publicado: (2024)
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
por: Yamamoto, Naoya, et al.
Publicado: (2025)
por: Yamamoto, Naoya, et al.
Publicado: (2025)
Inertial Newton Algorithms Avoiding Strict Saddle Points
por: Castera, Camille
Publicado: (2021)
por: Castera, Camille
Publicado: (2021)
Training-Free Vector Quantization via Gaussian VAEs
por: Xu, Tongda, et al.
Publicado: (2025)
por: Xu, Tongda, et al.
Publicado: (2025)
Neural Network-based High-index Saddle Dynamics Method for Searching Saddle Points and Solution Landscape
por: Liu, Yuankai, et al.
Publicado: (2024)
por: Liu, Yuankai, et al.
Publicado: (2024)
Better Training Data Attribution via Better Inverse Hessian-Vector Products
por: Wang, Andrew, et al.
Publicado: (2025)
por: Wang, Andrew, et al.
Publicado: (2025)
Getting Free Bits Back from Rotational Symmetries in LLMs
por: He, Jiajun, et al.
Publicado: (2024)
por: He, Jiajun, et al.
Publicado: (2024)
Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks
por: Corlouer, Guillaume, et al.
Publicado: (2026)
por: Corlouer, Guillaume, et al.
Publicado: (2026)
Uncertainty Modeling in Graph Neural Networks via Stochastic Differential Equations
por: Bergna, Richard, et al.
Publicado: (2024)
por: Bergna, Richard, et al.
Publicado: (2024)
Dimension-Free Saddle-Point Escape in Muon
por: Long, Yanlin, et al.
Publicado: (2026)
por: Long, Yanlin, et al.
Publicado: (2026)
RECOMBINER: Robust and Enhanced Compression with Bayesian Implicit Neural Representations
por: He, Jiajun, et al.
Publicado: (2023)
por: He, Jiajun, et al.
Publicado: (2023)
Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle Escape
por: Bantzis, Ioannis, et al.
Publicado: (2025)
por: Bantzis, Ioannis, et al.
Publicado: (2025)
Towards Quantifying the Hessian Structure of Neural Networks
por: Dong, Zhaorui, et al.
Publicado: (2025)
por: Dong, Zhaorui, et al.
Publicado: (2025)
Online Newton Method for Bandit Convex Optimisation
por: Fokkema, Hidde, et al.
Publicado: (2024)
por: Fokkema, Hidde, et al.
Publicado: (2024)
Characterizing Learning in Deep Neural Networks using Tractable Algorithmic Complexity Analysis
por: Bakhtiarifard, Pedram, et al.
Publicado: (2026)
por: Bakhtiarifard, Pedram, et al.
Publicado: (2026)
Training Neural Samplers with Reverse Diffusive KL Divergence
por: He, Jiajun, et al.
Publicado: (2024)
por: He, Jiajun, et al.
Publicado: (2024)
Newton-CG methods for nonconvex unconstrained optimization with Hölder continuous Hessian
por: He, Chuan, et al.
Publicado: (2023)
por: He, Chuan, et al.
Publicado: (2023)
A Gaussian Process View on Observation Noise and Initialization in Wide Neural Networks
por: Calvo-Ordoñez, Sergio, et al.
Publicado: (2025)
por: Calvo-Ordoñez, Sergio, et al.
Publicado: (2025)
Exact, Tractable Gauss-Newton Optimization in Deep Reversible Architectures Reveal Poor Generalization
por: Buffelli, Davide, et al.
Publicado: (2024)
por: Buffelli, Davide, et al.
Publicado: (2024)
Causal Effect Estimation under Networked Interference without Networked Unconfoundedness Assumption
por: Chen, Weilin, et al.
Publicado: (2025)
por: Chen, Weilin, et al.
Publicado: (2025)
SGD with Partial Hessian for Deep Neural Networks Optimization
por: Sun, Ying, et al.
Publicado: (2024)
por: Sun, Ying, et al.
Publicado: (2024)
Diagnosing and fixing common problems in Bayesian optimization for molecule design
por: Tripp, Austin, et al.
Publicado: (2024)
por: Tripp, Austin, et al.
Publicado: (2024)
Stein Variational Newton Neural Network Ensembles
por: Flöge, Klemens, et al.
Publicado: (2024)
por: Flöge, Klemens, et al.
Publicado: (2024)
From Saddle Points Toward Global Minima: A Newton-Type Method on Wasserstein Space
por: Lascu, Razvan-Andrei, et al.
Publicado: (2026)
por: Lascu, Razvan-Andrei, et al.
Publicado: (2026)
Geometry of Critical Sets and Existence of Saddle Branches for Two-layer Neural Networks
por: Zhang, Leyang, et al.
Publicado: (2024)
por: Zhang, Leyang, et al.
Publicado: (2024)
Directional Convergence Near Small Initializations and Saddles in Two-Homogeneous Neural Networks
por: Kumar, Akshay, et al.
Publicado: (2024)
por: Kumar, Akshay, et al.
Publicado: (2024)
Loss Landscape of Shallow ReLU-like Neural Networks: Stationary Points, Saddle Escape, and Network Embedding
por: Wu, Frank Zhengqing, et al.
Publicado: (2024)
por: Wu, Frank Zhengqing, et al.
Publicado: (2024)
On Newton's Method to Unlearn Neural Networks
por: Bui, Nhung, et al.
Publicado: (2024)
por: Bui, Nhung, et al.
Publicado: (2024)
Tractable Probabilistic Graph Representation Learning with Graph-Induced Sum-Product Networks
por: Errica, Federico, et al.
Publicado: (2023)
por: Errica, Federico, et al.
Publicado: (2023)
Sum-Product-Set Networks: Deep Tractable Models for Tree-Structured Graphs
por: Papež, Milan, et al.
Publicado: (2024)
por: Papež, Milan, et al.
Publicado: (2024)
Decoupled PFNs: Identifiable Epistemic-Aleatoric Decomposition via Structured Synthetic Priors
por: Bergna, Richard, et al.
Publicado: (2026)
por: Bergna, Richard, et al.
Publicado: (2026)
There Was Never a Bottleneck in Concept Bottleneck Models
por: Almudévar, Antonio, et al.
Publicado: (2025)
por: Almudévar, Antonio, et al.
Publicado: (2025)
Post-Hoc Uncertainty Quantification in Pre-Trained Neural Networks via Activation-Level Gaussian Processes
por: Bergna, Richard, et al.
Publicado: (2025)
por: Bergna, Richard, et al.
Publicado: (2025)
Theoretical characterisation of the Gauss-Newton conditioning in Neural Networks
por: Zhao, Jim, et al.
Publicado: (2024)
por: Zhao, Jim, et al.
Publicado: (2024)
Hessian-Free Online Certified Unlearning
por: Qiao, Xinbao, et al.
Publicado: (2024)
por: Qiao, Xinbao, et al.
Publicado: (2024)
Revisit, Extend, and Enhance Hessian-Free Influence Functions
por: Yang, Ziao, et al.
Publicado: (2024)
por: Yang, Ziao, et al.
Publicado: (2024)
Ejemplares similares
-
Studying K-FAC Heuristics by Viewing Adam through a Second-Order Lens
por: Clarke, Ross M., et al.
Publicado: (2023) -
Warm Start Marginal Likelihood Optimisation for Iterative Gaussian Processes
por: Lin, Jihao Andreas, et al.
Publicado: (2024) -
Leveraging Task Structures for Improved Identifiability in Neural Network Representations
por: Chen, Wenlin, et al.
Publicado: (2023) -
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
por: Zhang, Yedi, et al.
Publicado: (2025) -
Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes
por: Lin, Jihao Andreas, et al.
Publicado: (2024)