Studying K-FAC Heuristics by Viewing Adam through a Second-Order Lens
Fuente:
arXiv
Saved in:
| Main Authors: | Clarke, Ross M., Hernández-Lobato, José Miguel |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Series of Hessian-Vector Products for Tractable Saddle-Free Newton Optimisation of Neural Networks
by: Oldewage, Elre T., et al.
Published: (2023)
by: Oldewage, Elre T., et al.
Published: (2023)
From Adam to Adam-Like Lagrangians: Second-Order Nonlocal Dynamics
by: Heredia, Carlos
Published: (2026)
by: Heredia, Carlos
Published: (2026)
Aligning Multimodal Representations through an Information Bottleneck
by: Almudévar, Antonio, et al.
Published: (2025)
by: Almudévar, Antonio, et al.
Published: (2025)
Beyond SGD, Without SVD: Proximal Subspace Iteration LoRA with Diagonal Fractional K-FAC
by: Almansoori, Abdulla Jasem, et al.
Published: (2026)
by: Almansoori, Abdulla Jasem, et al.
Published: (2026)
There Was Never a Bottleneck in Concept Bottleneck Models
by: Almudévar, Antonio, et al.
Published: (2025)
by: Almudévar, Antonio, et al.
Published: (2025)
Diagnosing and fixing common problems in Bayesian optimization for molecule design
by: Tripp, Austin, et al.
Published: (2024)
by: Tripp, Austin, et al.
Published: (2024)
A Gaussian Process View on Observation Noise and Initialization in Wide Neural Networks
by: Calvo-Ordoñez, Sergio, et al.
Published: (2025)
by: Calvo-Ordoñez, Sergio, et al.
Published: (2025)
Decoupled PFNs: Identifiable Epistemic-Aleatoric Decomposition via Structured Synthetic Priors
by: Bergna, Richard, et al.
Published: (2026)
by: Bergna, Richard, et al.
Published: (2026)
Rethinking Adam for Time Series Forecasting: A Simple Heuristic to Improve Optimization under Distribution Shifts
by: Dong, Yuze, et al.
Published: (2026)
by: Dong, Yuze, et al.
Published: (2026)
Online Laplace Model Selection Revisited
by: Lin, Jihao Andreas, et al.
Published: (2023)
by: Lin, Jihao Andreas, et al.
Published: (2023)
Conditional Diffusion Sampling
by: Castro-Macías, Francisco M., et al.
Published: (2026)
by: Castro-Macías, Francisco M., et al.
Published: (2026)
Efficient and Unbiased Sampling from Boltzmann Distributions via Variance-Tuned Diffusion Models
by: Zhang, Fengzhe, et al.
Published: (2025)
by: Zhang, Fengzhe, et al.
Published: (2025)
Accelerating Relative Entropy Coding with Space Partitioning
by: He, Jiajun, et al.
Published: (2024)
by: He, Jiajun, et al.
Published: (2024)
Getting Free Bits Back from Rotational Symmetries in LLMs
by: He, Jiajun, et al.
Published: (2024)
by: He, Jiajun, et al.
Published: (2024)
RECOMBINER: Robust and Enhanced Compression with Bayesian Implicit Neural Representations
by: He, Jiajun, et al.
Published: (2023)
by: He, Jiajun, et al.
Published: (2023)
Leveraging Task Structures for Improved Identifiability in Neural Network Representations
by: Chen, Wenlin, et al.
Published: (2023)
by: Chen, Wenlin, et al.
Published: (2023)
Improving Iterative Gaussian Processes via Warm Starting Sequential Posteriors
by: Dong, Alan Yufei, et al.
Published: (2025)
by: Dong, Alan Yufei, et al.
Published: (2025)
Mitigating Forgetting in Low Rank Adaptation
by: Sliwa, Joanna, et al.
Published: (2025)
by: Sliwa, Joanna, et al.
Published: (2025)
A Diffusive Classification Loss for Learning Energy-based Generative Models
by: OuYang, RuiKang, et al.
Published: (2026)
by: OuYang, RuiKang, et al.
Published: (2026)
RNE: plug-and-play diffusion inference-time control and energy-based training
by: He, Jiajun, et al.
Published: (2025)
by: He, Jiajun, et al.
Published: (2025)
SL-FAC: A Communication-Efficient Split Learning Framework with Frequency-Aware Compression
by: Lin, Zehang, et al.
Published: (2026)
by: Lin, Zehang, et al.
Published: (2026)
Refresh-Scaling the Memory of Balanced Adam
by: Fernández-Hernández, Alberto, et al.
Published: (2026)
by: Fernández-Hernández, Alberto, et al.
Published: (2026)
Calibration through the Lens of Interpretability
by: Torabian, Alireza, et al.
Published: (2024)
by: Torabian, Alireza, et al.
Published: (2024)
Fast Second-Order Online Kernel Learning through Incremental Matrix Sketching and Decomposition
by: Wen, Dongxie, et al.
Published: (2024)
by: Wen, Dongxie, et al.
Published: (2024)
Warm Start Marginal Likelihood Optimisation for Iterative Gaussian Processes
by: Lin, Jihao Andreas, et al.
Published: (2024)
by: Lin, Jihao Andreas, et al.
Published: (2024)
The Minimax Rate of Second-Order Calibration
by: Ciosek, Kamil, et al.
Published: (2026)
by: Ciosek, Kamil, et al.
Published: (2026)
LossLens: Diagnostics for Machine Learning through Loss Landscape Visual Analytics
by: Xie, Tiankai, et al.
Published: (2024)
by: Xie, Tiankai, et al.
Published: (2024)
One Size Fits None: Heuristic Collapse in LLM Investment Advice
by: Ross, Jillian, et al.
Published: (2026)
by: Ross, Jillian, et al.
Published: (2026)
Nonparametric Heterogeneous Long-term Causal Effect Estimation via Data Combination
by: Chen, Weilin, et al.
Published: (2025)
by: Chen, Weilin, et al.
Published: (2025)
Training Neural Samplers with Reverse Diffusive KL Divergence
by: He, Jiajun, et al.
Published: (2024)
by: He, Jiajun, et al.
Published: (2024)
Causal Effect Estimation under Networked Interference without Networked Unconfoundedness Assumption
by: Chen, Weilin, et al.
Published: (2025)
by: Chen, Weilin, et al.
Published: (2025)
Long-term Causal Inference via Modeling Sequential Latent Confounding
by: Chen, Weilin, et al.
Published: (2025)
by: Chen, Weilin, et al.
Published: (2025)
FOSI: Hybrid First and Second Order Optimization
by: Sivan, Hadar, et al.
Published: (2023)
by: Sivan, Hadar, et al.
Published: (2023)
The Power of Second Order Methods for Sequence Preconditioning
by: Marsden, Annie, et al.
Published: (2026)
by: Marsden, Annie, et al.
Published: (2026)
Second Order Methods for Bandit Optimization and Control
by: Suggala, Arun, et al.
Published: (2024)
by: Suggala, Arun, et al.
Published: (2024)
Deep Generative Models through the Lens of the Manifold Hypothesis: A Survey and New Connections
by: Loaiza-Ganem, Gabriel, et al.
Published: (2024)
by: Loaiza-Ganem, Gabriel, et al.
Published: (2024)
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
by: Jin, Ruinan, et al.
Published: (2026)
by: Jin, Ruinan, et al.
Published: (2026)
Mitigating Task-Order Sensitivity and Forgetting via Hierarchical Second-Order Consolidation
by: Nag, Protik, et al.
Published: (2026)
by: Nag, Protik, et al.
Published: (2026)
In Search of Adam's Secret Sauce
by: Orvieto, Antonio, et al.
Published: (2025)
by: Orvieto, Antonio, et al.
Published: (2025)
On Equivariant Model Selection through the Lens of Uncertainty
by: van der Linden, Putri A., et al.
Published: (2025)
by: van der Linden, Putri A., et al.
Published: (2025)
Similar Items
-
Series of Hessian-Vector Products for Tractable Saddle-Free Newton Optimisation of Neural Networks
by: Oldewage, Elre T., et al.
Published: (2023) -
From Adam to Adam-Like Lagrangians: Second-Order Nonlocal Dynamics
by: Heredia, Carlos
Published: (2026) -
Aligning Multimodal Representations through an Information Bottleneck
by: Almudévar, Antonio, et al.
Published: (2025) -
Beyond SGD, Without SVD: Proximal Subspace Iteration LoRA with Diagonal Fractional K-FAC
by: Almansoori, Abdulla Jasem, et al.
Published: (2026) -
There Was Never a Bottleneck in Concept Bottleneck Models
by: Almudévar, Antonio, et al.
Published: (2025)