Position: Curvature Matrices Should Be Democratized via Linear Operators
Fuente:
arXiv
Guardado en:
| Autores principales: | Dangel, Felix, Eschenhagen, Runa, Ormaniec, Weronika, Fernandez, Andres, Tatzel, Lukas, Kristiadi, Agustinus |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On the Disconnect Between Theory and Practice of Neural Networks: Limits of the NTK Perspective
por: Wenger, Jonathan, et al.
Publicado: (2023)
por: Wenger, Jonathan, et al.
Publicado: (2023)
Structured Inverse-Free Natural Gradient: Memory-Efficient & Numerically-Stable KFAC
por: Lin, Wu, et al.
Publicado: (2023)
por: Lin, Wu, et al.
Publicado: (2023)
Kronecker-factored Approximate Curvature (KFAC) From Scratch
por: Dangel, Felix, et al.
Publicado: (2025)
por: Dangel, Felix, et al.
Publicado: (2025)
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
por: Ormaniec, Weronika, et al.
Publicado: (2024)
por: Ormaniec, Weronika, et al.
Publicado: (2024)
Spectral-factorized Positive-definite Curvature Learning for NN Training
por: Lin, Wu, et al.
Publicado: (2025)
por: Lin, Wu, et al.
Publicado: (2025)
Introduction to the Analysis of Probabilistic Decision-Making Algorithms
por: Kristiadi, Agustinus
Publicado: (2025)
por: Kristiadi, Agustinus
Publicado: (2025)
Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization
por: Lin, Wu, et al.
Publicado: (2025)
por: Lin, Wu, et al.
Publicado: (2025)
Fusion of Graph Neural Networks via Optimal Transport
por: Ormaniec, Weronika, et al.
Publicado: (2025)
por: Ormaniec, Weronika, et al.
Publicado: (2025)
Better Hessians Matter: Studying the Impact of Curvature Approximations in Influence Functions
por: Hong, Steve, et al.
Publicado: (2025)
por: Hong, Steve, et al.
Publicado: (2025)
Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective
por: Lin, Wu, et al.
Publicado: (2024)
por: Lin, Wu, et al.
Publicado: (2024)
Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
por: Cinquin, Tristan, et al.
Publicado: (2025)
por: Cinquin, Tristan, et al.
Publicado: (2025)
Sketching Low-Rank Plus Diagonal Matrices
por: Fernandez, Andres, et al.
Publicado: (2025)
por: Fernandez, Andres, et al.
Publicado: (2025)
Kronecker-Factored Approximate Curvature for Modern Neural Network Architectures
por: Eschenhagen, Runa, et al.
Publicado: (2023)
por: Eschenhagen, Runa, et al.
Publicado: (2023)
FlashMD: long-stride, universal prediction of molecular dynamics
por: Bigi, Filippo, et al.
Publicado: (2025)
por: Bigi, Filippo, et al.
Publicado: (2025)
Low-Rank Filtering and Smoothing for Sequential Deep Learning
por: Sliwa, Joanna, et al.
Publicado: (2024)
por: Sliwa, Joanna, et al.
Publicado: (2024)
Preventing Arbitrarily High Confidence on Far-Away Data in Point-Estimated Discriminative Neural Networks
por: Rashid, Ahmad, et al.
Publicado: (2023)
por: Rashid, Ahmad, et al.
Publicado: (2023)
Kronecker-Factored Approximate Curvature for Physics-Informed Neural Networks
por: Dangel, Felix, et al.
Publicado: (2024)
por: Dangel, Felix, et al.
Publicado: (2024)
A Critical Look At Tokenwise Reward-Guided Text Generation
por: Rashid, Ahmad, et al.
Publicado: (2024)
por: Rashid, Ahmad, et al.
Publicado: (2024)
Standardizing Structural Causal Models
por: Ormaniec, Weronika, et al.
Publicado: (2024)
por: Ormaniec, Weronika, et al.
Publicado: (2024)
A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?
por: Kristiadi, Agustinus, et al.
Publicado: (2024)
por: Kristiadi, Agustinus, et al.
Publicado: (2024)
How Useful is Intermittent, Asynchronous Expert Feedback for Bayesian Optimization?
por: Kristiadi, Agustinus, et al.
Publicado: (2024)
por: Kristiadi, Agustinus, et al.
Publicado: (2024)
Convolutions and More as Einsum: A Tensor Network Perspective with Advances for Second-Order Methods
por: Dangel, Felix
Publicado: (2023)
por: Dangel, Felix
Publicado: (2023)
Towards Cost-Effective Reward Guided Text Generation
por: Rashid, Ahmad, et al.
Publicado: (2025)
por: Rashid, Ahmad, et al.
Publicado: (2025)
Lowering PyTorch's Memory Consumption for Selective Differentiation
por: Bhatia, Samarth, et al.
Publicado: (2024)
por: Bhatia, Samarth, et al.
Publicado: (2024)
Uncertainty-Guided Likelihood Tree Search
por: Grosse, Julia, et al.
Publicado: (2024)
por: Grosse, Julia, et al.
Publicado: (2024)
Accelerating Non-Conjugate Gaussian Processes By Trading Off Computation For Uncertainty
por: Tatzel, Lukas, et al.
Publicado: (2023)
por: Tatzel, Lukas, et al.
Publicado: (2023)
Debiasing Mini-Batch Quadratics for Applications in Deep Learning
por: Tatzel, Lukas, et al.
Publicado: (2024)
por: Tatzel, Lukas, et al.
Publicado: (2024)
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
por: Eschenhagen, Runa, et al.
Publicado: (2026)
por: Eschenhagen, Runa, et al.
Publicado: (2026)
Improving Energy Natural Gradient Descent through Woodbury, Momentum, and Randomization
por: Guzmán-Cordero, Andrés, et al.
Publicado: (2025)
por: Guzmán-Cordero, Andrés, et al.
Publicado: (2025)
Efficient Bilevel Optimization with KFAC-Based Hypergradients
por: Liao, Disen, et al.
Publicado: (2026)
por: Liao, Disen, et al.
Publicado: (2026)
Influence Functions for Scalable Data Attribution in Diffusion Models
por: Mlodozeniec, Bruno, et al.
Publicado: (2024)
por: Mlodozeniec, Bruno, et al.
Publicado: (2024)
One Set to Rule Them All: How to Obtain General Chemical Conditions via Bayesian Optimization over Curried Functions
por: Schmid, Stefan P., et al.
Publicado: (2025)
por: Schmid, Stefan P., et al.
Publicado: (2025)
Transition Constrained Bayesian Optimization via Markov Decision Processes
por: Folch, Jose Pablo, et al.
Publicado: (2024)
por: Folch, Jose Pablo, et al.
Publicado: (2024)
Protein Counterfactuals via Diffusion-Guided Latent Optimization
por: Kłos, Weronika, et al.
Publicado: (2026)
por: Kłos, Weronika, et al.
Publicado: (2026)
Hide & Seek: Transformer Symmetries Obscure Sharpness & Riemannian Geometry Finds It
por: da Silva, Marvin F., et al.
Publicado: (2025)
por: da Silva, Marvin F., et al.
Publicado: (2025)
Collapsing Taylor Mode Automatic Differentiation
por: Dangel, Felix, et al.
Publicado: (2025)
por: Dangel, Felix, et al.
Publicado: (2025)
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
por: Eschenhagen, Runa, et al.
Publicado: (2025)
por: Eschenhagen, Runa, et al.
Publicado: (2025)
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
por: Li, YuXin, et al.
Publicado: (2025)
por: Li, YuXin, et al.
Publicado: (2025)
Effective Structural Encodings via Local Curvature Profiles
por: Fesser, Lukas, et al.
Publicado: (2023)
por: Fesser, Lukas, et al.
Publicado: (2023)
Generalizing the Geometry of Model Merging Through Frechet Averages
por: da Silva, Marvin F., et al.
Publicado: (2026)
por: da Silva, Marvin F., et al.
Publicado: (2026)
Ejemplares similares
-
On the Disconnect Between Theory and Practice of Neural Networks: Limits of the NTK Perspective
por: Wenger, Jonathan, et al.
Publicado: (2023) -
Structured Inverse-Free Natural Gradient: Memory-Efficient & Numerically-Stable KFAC
por: Lin, Wu, et al.
Publicado: (2023) -
Kronecker-factored Approximate Curvature (KFAC) From Scratch
por: Dangel, Felix, et al.
Publicado: (2025) -
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
por: Ormaniec, Weronika, et al.
Publicado: (2024) -
Spectral-factorized Positive-definite Curvature Learning for NN Training
por: Lin, Wu, et al.
Publicado: (2025)