Spectral-factorized Positive-definite Curvature Learning for NN Training
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Wu, Dangel, Felix, Eschenhagen, Runa, Bae, Juhan, Turner, Richard E., Grosse, Roger B. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective
by: Lin, Wu, et al.
Published: (2024)
by: Lin, Wu, et al.
Published: (2024)
Kronecker-factored Approximate Curvature (KFAC) From Scratch
by: Dangel, Felix, et al.
Published: (2025)
by: Dangel, Felix, et al.
Published: (2025)
Position: Curvature Matrices Should Be Democratized via Linear Operators
by: Dangel, Felix, et al.
Published: (2025)
by: Dangel, Felix, et al.
Published: (2025)
Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization
by: Lin, Wu, et al.
Published: (2025)
by: Lin, Wu, et al.
Published: (2025)
Structured Inverse-Free Natural Gradient: Memory-Efficient & Numerically-Stable KFAC
by: Lin, Wu, et al.
Published: (2023)
by: Lin, Wu, et al.
Published: (2023)
Better Hessians Matter: Studying the Impact of Curvature Approximations in Influence Functions
by: Hong, Steve, et al.
Published: (2025)
by: Hong, Steve, et al.
Published: (2025)
Influence Functions for Scalable Data Attribution in Diffusion Models
by: Mlodozeniec, Bruno, et al.
Published: (2024)
by: Mlodozeniec, Bruno, et al.
Published: (2024)
Kronecker-Factored Approximate Curvature for Modern Neural Network Architectures
by: Eschenhagen, Runa, et al.
Published: (2023)
by: Eschenhagen, Runa, et al.
Published: (2023)
Training Data Attribution via Approximate Unrolled Differentiation
by: Bae, Juhan, et al.
Published: (2024)
by: Bae, Juhan, et al.
Published: (2024)
Kronecker-Factored Approximate Curvature for Physics-Informed Neural Networks
by: Dangel, Felix, et al.
Published: (2024)
by: Dangel, Felix, et al.
Published: (2024)
Better Training Data Attribution via Better Inverse Hessian-Vector Products
by: Wang, Andrew, et al.
Published: (2025)
by: Wang, Andrew, et al.
Published: (2025)
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
by: Eschenhagen, Runa, et al.
Published: (2026)
by: Eschenhagen, Runa, et al.
Published: (2026)
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
by: Eschenhagen, Runa, et al.
Published: (2025)
by: Eschenhagen, Runa, et al.
Published: (2025)
Accelerating Neural Network Training: An Analysis of the AlgoPerf Competition
by: Kasimbeg, Priya, et al.
Published: (2025)
by: Kasimbeg, Priya, et al.
Published: (2025)
Convolutions and More as Einsum: A Tensor Network Perspective with Advances for Second-Order Methods
by: Dangel, Felix
Published: (2023)
by: Dangel, Felix
Published: (2023)
Lowering PyTorch's Memory Consumption for Selective Differentiation
by: Bhatia, Samarth, et al.
Published: (2024)
by: Bhatia, Samarth, et al.
Published: (2024)
Exploring Training Data Attribution under Limited Access Constraints
by: Zhang, Shiyuan, et al.
Published: (2025)
by: Zhang, Shiyuan, et al.
Published: (2025)
On the Disconnect Between Theory and Practice of Neural Networks: Limits of the NTK Perspective
by: Wenger, Jonathan, et al.
Published: (2023)
by: Wenger, Jonathan, et al.
Published: (2023)
Efficient Bilevel Optimization with KFAC-Based Hypergradients
by: Liao, Disen, et al.
Published: (2026)
by: Liao, Disen, et al.
Published: (2026)
Gauss-Newton Unlearning for the LLM Era
by: McKinney, Lev, et al.
Published: (2026)
by: McKinney, Lev, et al.
Published: (2026)
Distributional Training Data Attribution: What do Influence Functions Sample?
by: Mlodozeniec, Bruno, et al.
Published: (2025)
by: Mlodozeniec, Bruno, et al.
Published: (2025)
Simplifying Momentum-based Positive-definite Submanifold Optimization with Applications to Deep Learning
by: Lin, Wu, et al.
Published: (2023)
by: Lin, Wu, et al.
Published: (2023)
Reparametrizing Shampoo and SOAP for Subspace Basis Updates and BFloat16 Storage
by: Milligan, Alan, et al.
Published: (2026)
by: Milligan, Alan, et al.
Published: (2026)
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
by: Ormaniec, Weronika, et al.
Published: (2024)
by: Ormaniec, Weronika, et al.
Published: (2024)
Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement Learning
by: Elsayed, Mohamed, et al.
Published: (2024)
by: Elsayed, Mohamed, et al.
Published: (2024)
Hide & Seek: Transformer Symmetries Obscure Sharpness & Riemannian Geometry Finds It
by: da Silva, Marvin F., et al.
Published: (2025)
by: da Silva, Marvin F., et al.
Published: (2025)
Collapsing Taylor Mode Automatic Differentiation
by: Dangel, Felix, et al.
Published: (2025)
by: Dangel, Felix, et al.
Published: (2025)
Benchmarking Neural Network Training Algorithms
by: Dahl, George E., et al.
Published: (2023)
by: Dahl, George E., et al.
Published: (2023)
Improving Energy Natural Gradient Descent through Woodbury, Momentum, and Randomization
by: Guzmán-Cordero, Andrés, et al.
Published: (2025)
by: Guzmán-Cordero, Andrés, et al.
Published: (2025)
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
by: Li, YuXin, et al.
Published: (2025)
by: Li, YuXin, et al.
Published: (2025)
Sketching Low-Rank Plus Diagonal Matrices
by: Fernandez, Andres, et al.
Published: (2025)
by: Fernandez, Andres, et al.
Published: (2025)
IF-GUIDE: Influence Function-Guided Detoxification of LLMs
by: Coalson, Zachary, et al.
Published: (2025)
by: Coalson, Zachary, et al.
Published: (2025)
Generalizing the Geometry of Model Merging Through Frechet Averages
by: da Silva, Marvin F., et al.
Published: (2026)
by: da Silva, Marvin F., et al.
Published: (2026)
UnifiedNN: Efficient Neural Network Training on the Cloud
by: Taki, Sifat Ut, et al.
Published: (2024)
by: Taki, Sifat Ut, et al.
Published: (2024)
REFACTOR: Learning to Extract Theorems from Proofs
by: Zhou, Jin Peng, et al.
Published: (2024)
by: Zhou, Jin Peng, et al.
Published: (2024)
Geometric Embedding Alignment via Curvature Matching in Transfer Learning
by: Ko, Sung Moon, et al.
Published: (2025)
by: Ko, Sung Moon, et al.
Published: (2025)
Using Large Language Models for Hyperparameter Optimization
by: Zhang, Michael R., et al.
Published: (2023)
by: Zhang, Michael R., et al.
Published: (2023)
An Introduction to Transformers
by: Turner, Richard E.
Published: (2023)
by: Turner, Richard E.
Published: (2023)
Newton Losses: Using Curvature Information for Learning with Differentiable Algorithms
by: Petersen, Felix, et al.
Published: (2024)
by: Petersen, Felix, et al.
Published: (2024)
CAT: Curvature-Adaptive Transformers for Geometry-Aware Learning
by: Lin, Ryan Y., et al.
Published: (2025)
by: Lin, Ryan Y., et al.
Published: (2025)
Similar Items
-
Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective
by: Lin, Wu, et al.
Published: (2024) -
Kronecker-factored Approximate Curvature (KFAC) From Scratch
by: Dangel, Felix, et al.
Published: (2025) -
Position: Curvature Matrices Should Be Democratized via Linear Operators
by: Dangel, Felix, et al.
Published: (2025) -
Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization
by: Lin, Wu, et al.
Published: (2025) -
Structured Inverse-Free Natural Gradient: Memory-Efficient & Numerically-Stable KFAC
by: Lin, Wu, et al.
Published: (2023)