Structured Inverse-Free Natural Gradient: Memory-Efficient & Numerically-Stable KFAC
Fuente:
arXiv
Salvato in:
| Autori principali: | Lin, Wu, Dangel, Felix, Eschenhagen, Runa, Neklyudov, Kirill, Kristiadi, Agustinus, Turner, Richard E., Makhzani, Alireza |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective
di: Lin, Wu, et al.
Pubblicazione: (2024)
di: Lin, Wu, et al.
Pubblicazione: (2024)
Kronecker-factored Approximate Curvature (KFAC) From Scratch
di: Dangel, Felix, et al.
Pubblicazione: (2025)
di: Dangel, Felix, et al.
Pubblicazione: (2025)
Position: Curvature Matrices Should Be Democratized via Linear Operators
di: Dangel, Felix, et al.
Pubblicazione: (2025)
di: Dangel, Felix, et al.
Pubblicazione: (2025)
On the Disconnect Between Theory and Practice of Neural Networks: Limits of the NTK Perspective
di: Wenger, Jonathan, et al.
Pubblicazione: (2023)
di: Wenger, Jonathan, et al.
Pubblicazione: (2023)
Efficient Bilevel Optimization with KFAC-Based Hypergradients
di: Liao, Disen, et al.
Pubblicazione: (2026)
di: Liao, Disen, et al.
Pubblicazione: (2026)
Spectral-factorized Positive-definite Curvature Learning for NN Training
di: Lin, Wu, et al.
Pubblicazione: (2025)
di: Lin, Wu, et al.
Pubblicazione: (2025)
Introduction to the Analysis of Probabilistic Decision-Making Algorithms
di: Kristiadi, Agustinus
Pubblicazione: (2025)
di: Kristiadi, Agustinus
Pubblicazione: (2025)
A Computational Framework for Solving Wasserstein Lagrangian Flows
di: Neklyudov, Kirill, et al.
Pubblicazione: (2023)
di: Neklyudov, Kirill, et al.
Pubblicazione: (2023)
Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization
di: Lin, Wu, et al.
Pubblicazione: (2025)
di: Lin, Wu, et al.
Pubblicazione: (2025)
Better Hessians Matter: Studying the Impact of Curvature Approximations in Influence Functions
di: Hong, Steve, et al.
Pubblicazione: (2025)
di: Hong, Steve, et al.
Pubblicazione: (2025)
Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
di: Cinquin, Tristan, et al.
Pubblicazione: (2025)
di: Cinquin, Tristan, et al.
Pubblicazione: (2025)
Quantum HyperNetworks: Training Binary Neural Networks in Quantum Superposition
di: Carrasquilla, Juan, et al.
Pubblicazione: (2023)
di: Carrasquilla, Juan, et al.
Pubblicazione: (2023)
Kronecker-Factored Approximate Curvature for Modern Neural Network Architectures
di: Eschenhagen, Runa, et al.
Pubblicazione: (2023)
di: Eschenhagen, Runa, et al.
Pubblicazione: (2023)
FlashMD: long-stride, universal prediction of molecular dynamics
di: Bigi, Filippo, et al.
Pubblicazione: (2025)
di: Bigi, Filippo, et al.
Pubblicazione: (2025)
A Critical Look At Tokenwise Reward-Guided Text Generation
di: Rashid, Ahmad, et al.
Pubblicazione: (2024)
di: Rashid, Ahmad, et al.
Pubblicazione: (2024)
Influence Functions for Scalable Data Attribution in Diffusion Models
di: Mlodozeniec, Bruno, et al.
Pubblicazione: (2024)
di: Mlodozeniec, Bruno, et al.
Pubblicazione: (2024)
Lowering PyTorch's Memory Consumption for Selective Differentiation
di: Bhatia, Samarth, et al.
Pubblicazione: (2024)
di: Bhatia, Samarth, et al.
Pubblicazione: (2024)
Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner
di: Eschenhagen, Runa, et al.
Pubblicazione: (2025)
di: Eschenhagen, Runa, et al.
Pubblicazione: (2025)
Low-Rank Filtering and Smoothing for Sequential Deep Learning
di: Sliwa, Joanna, et al.
Pubblicazione: (2024)
di: Sliwa, Joanna, et al.
Pubblicazione: (2024)
Preventing Arbitrarily High Confidence on Far-Away Data in Point-Estimated Discriminative Neural Networks
di: Rashid, Ahmad, et al.
Pubblicazione: (2023)
di: Rashid, Ahmad, et al.
Pubblicazione: (2023)
Improving Energy Natural Gradient Descent through Woodbury, Momentum, and Randomization
di: Guzmán-Cordero, Andrés, et al.
Pubblicazione: (2025)
di: Guzmán-Cordero, Andrés, et al.
Pubblicazione: (2025)
Convolutions and More as Einsum: A Tensor Network Perspective with Advances for Second-Order Methods
di: Dangel, Felix
Pubblicazione: (2023)
di: Dangel, Felix
Pubblicazione: (2023)
Towards Cost-Effective Reward Guided Text Generation
di: Rashid, Ahmad, et al.
Pubblicazione: (2025)
di: Rashid, Ahmad, et al.
Pubblicazione: (2025)
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
di: Li, YuXin, et al.
Pubblicazione: (2025)
di: Li, YuXin, et al.
Pubblicazione: (2025)
Random Cycle Coding: Lossless Compression of Cluster Assignments via Bits-Back Coding
di: Severo, Daniel, et al.
Pubblicazione: (2024)
di: Severo, Daniel, et al.
Pubblicazione: (2024)
A Sober Look at LLMs for Material Discovery: Are They Actually Good for Bayesian Optimization Over Molecules?
di: Kristiadi, Agustinus, et al.
Pubblicazione: (2024)
di: Kristiadi, Agustinus, et al.
Pubblicazione: (2024)
How Useful is Intermittent, Asynchronous Expert Feedback for Bayesian Optimization?
di: Kristiadi, Agustinus, et al.
Pubblicazione: (2024)
di: Kristiadi, Agustinus, et al.
Pubblicazione: (2024)
Optimising Distributions with Natural Gradient Surrogates
di: So, Jonathan, et al.
Pubblicazione: (2023)
di: So, Jonathan, et al.
Pubblicazione: (2023)
A Call to Lagrangian Action: Learning Population Mechanics from Temporal Snapshots
di: Guan, Vincent, et al.
Pubblicazione: (2026)
di: Guan, Vincent, et al.
Pubblicazione: (2026)
Uncertainty-Guided Likelihood Tree Search
di: Grosse, Julia, et al.
Pubblicazione: (2024)
di: Grosse, Julia, et al.
Pubblicazione: (2024)
Probabilistic Inference in Language Models via Twisted Sequential Monte Carlo
di: Zhao, Stephen, et al.
Pubblicazione: (2024)
di: Zhao, Stephen, et al.
Pubblicazione: (2024)
Wavefunction Flows: Efficient Quantum Simulation of Continuous Flow Models
di: Layden, David, et al.
Pubblicazione: (2025)
di: Layden, David, et al.
Pubblicazione: (2025)
Clarifying Shampoo: Adapting Spectral Descent to Stochasticity and the Parameter Trajectory
di: Eschenhagen, Runa, et al.
Pubblicazione: (2026)
di: Eschenhagen, Runa, et al.
Pubblicazione: (2026)
Local Governments Accountability: A Content Analysis of the Financial Audit Reports
di: Agustinus Salle
Pubblicazione: (2020)
di: Agustinus Salle
Pubblicazione: (2020)
Inverse-Free Fast Natural Gradient Descent Method for Deep Learning
di: Ou, Xinwei, et al.
Pubblicazione: (2024)
di: Ou, Xinwei, et al.
Pubblicazione: (2024)
LLM Processes: Numerical Predictive Distributions Conditioned on Natural Language
di: Requeima, James, et al.
Pubblicazione: (2024)
di: Requeima, James, et al.
Pubblicazione: (2024)
Kronecker-Factored Approximate Curvature for Physics-Informed Neural Networks
di: Dangel, Felix, et al.
Pubblicazione: (2024)
di: Dangel, Felix, et al.
Pubblicazione: (2024)
Self-Refining Training for Amortized Density Functional Theory
di: Hassan, Majdi, et al.
Pubblicazione: (2025)
di: Hassan, Majdi, et al.
Pubblicazione: (2025)
Riemannian MeanFlow
di: Woo, Dongyeop, et al.
Pubblicazione: (2026)
di: Woo, Dongyeop, et al.
Pubblicazione: (2026)
Reparametrizing Shampoo and SOAP for Subspace Basis Updates and BFloat16 Storage
di: Milligan, Alan, et al.
Pubblicazione: (2026)
di: Milligan, Alan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective
di: Lin, Wu, et al.
Pubblicazione: (2024) -
Kronecker-factored Approximate Curvature (KFAC) From Scratch
di: Dangel, Felix, et al.
Pubblicazione: (2025) -
Position: Curvature Matrices Should Be Democratized via Linear Operators
di: Dangel, Felix, et al.
Pubblicazione: (2025) -
On the Disconnect Between Theory and Practice of Neural Networks: Limits of the NTK Perspective
di: Wenger, Jonathan, et al.
Pubblicazione: (2023) -
Efficient Bilevel Optimization with KFAC-Based Hypergradients
di: Liao, Disen, et al.
Pubblicazione: (2026)