Efficient Bilevel Optimization with KFAC-Based Hypergradients
Fuente:
arXiv
Saved in:
| Main Authors: | Liao, Disen, Dangel, Felix, Yu, Yaoliang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Kronecker-factored Approximate Curvature (KFAC) From Scratch
by: Dangel, Felix, et al.
Published: (2025)
by: Dangel, Felix, et al.
Published: (2025)
Structured Inverse-Free Natural Gradient: Memory-Efficient & Numerically-Stable KFAC
by: Lin, Wu, et al.
Published: (2023)
by: Lin, Wu, et al.
Published: (2023)
Efficient Curvature-Aware Hypergradient Approximation for Bilevel Optimization
by: Dong, Youran, et al.
Published: (2025)
by: Dong, Youran, et al.
Published: (2025)
Convolutions and More as Einsum: A Tensor Network Perspective with Advances for Second-Order Methods
by: Dangel, Felix
Published: (2023)
by: Dangel, Felix
Published: (2023)
One Sample Fits All: Approximating All Probabilistic Values Simultaneously and Efficiently
by: Li, Weida, et al.
Published: (2024)
by: Li, Weida, et al.
Published: (2024)
Lowering PyTorch's Memory Consumption for Selective Differentiation
by: Bhatia, Samarth, et al.
Published: (2024)
by: Bhatia, Samarth, et al.
Published: (2024)
Convergence Properties of Stochastic Hypergradients
by: Grazzi, Riccardo, et al.
Published: (2020)
by: Grazzi, Riccardo, et al.
Published: (2020)
SFBD Flow: A Continuous-Optimization Framework for Training Diffusion Models with Noisy Samples
by: Lu, Haoye, et al.
Published: (2025)
by: Lu, Haoye, et al.
Published: (2025)
Glocal Hypergradient Estimation with Koopman Operator
by: Hataya, Ryuichiro, et al.
Published: (2024)
by: Hataya, Ryuichiro, et al.
Published: (2024)
Bi-Level Policy Optimization with Nyström Hypergradients
by: Prakash, Arjun, et al.
Published: (2025)
by: Prakash, Arjun, et al.
Published: (2025)
On the Disconnect Between Theory and Practice of Neural Networks: Limits of the NTK Perspective
by: Wenger, Jonathan, et al.
Published: (2023)
by: Wenger, Jonathan, et al.
Published: (2023)
HYDRA: Hypergradient Data Relevance Analysis for Interpreting Deep Neural Networks
by: Chen, Yuanyuan, et al.
Published: (2021)
by: Chen, Yuanyuan, et al.
Published: (2021)
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
by: Ormaniec, Weronika, et al.
Published: (2024)
by: Ormaniec, Weronika, et al.
Published: (2024)
Enhancing Hypergradients Estimation: A Study of Preconditioning and Reparameterization
by: Ye, Zhenzhang, et al.
Published: (2024)
by: Ye, Zhenzhang, et al.
Published: (2024)
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
by: Li, YuXin, et al.
Published: (2025)
by: Li, YuXin, et al.
Published: (2025)
Kronecker-Factored Approximate Curvature for Physics-Informed Neural Networks
by: Dangel, Felix, et al.
Published: (2024)
by: Dangel, Felix, et al.
Published: (2024)
Semi-Variance Reduction for Fair Federated Learning
by: Malekmohammadi, Saber, et al.
Published: (2024)
by: Malekmohammadi, Saber, et al.
Published: (2024)
A Unified Framework for Gradient Aggregation in Multi-Objective Optimization
by: Hu, Zeou, et al.
Published: (2026)
by: Hu, Zeou, et al.
Published: (2026)
Hide & Seek: Transformer Symmetries Obscure Sharpness & Riemannian Geometry Finds It
by: da Silva, Marvin F., et al.
Published: (2025)
by: da Silva, Marvin F., et al.
Published: (2025)
Collapsing Taylor Mode Automatic Differentiation
by: Dangel, Felix, et al.
Published: (2025)
by: Dangel, Felix, et al.
Published: (2025)
Stochastic Forward-Backward Deconvolution: Training Diffusion Models with Finite Noisy Datasets
by: Lu, Haoye, et al.
Published: (2025)
by: Lu, Haoye, et al.
Published: (2025)
SFBD-OMNI: Bridge models for lossy measurement restoration with limited clean samples
by: Lu, Haoye, et al.
Published: (2025)
by: Lu, Haoye, et al.
Published: (2025)
Adaptive Context Length Optimization with Low-Frequency Truncation for Multi-Agent Reinforcement Learning
by: Duan, Wenchang, et al.
Published: (2025)
by: Duan, Wenchang, et al.
Published: (2025)
Fairness-informed Pareto Optimization : An Efficient Bilevel Framework
by: Tanji, Sofiane, et al.
Published: (2026)
by: Tanji, Sofiane, et al.
Published: (2026)
DiffBreak: Is Diffusion-Based Purification Robust?
by: Kassis, Andre, et al.
Published: (2024)
by: Kassis, Andre, et al.
Published: (2024)
Improving Energy Natural Gradient Descent through Woodbury, Momentum, and Randomization
by: Guzmán-Cordero, Andrés, et al.
Published: (2025)
by: Guzmán-Cordero, Andrés, et al.
Published: (2025)
Federated Learning with Hypergradient-based Online Update of Aggregation Weights
by: Nakai-Kasai, Ayano, et al.
Published: (2026)
by: Nakai-Kasai, Ayano, et al.
Published: (2026)
Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent
by: Chu, Ya-Chi, et al.
Published: (2025)
by: Chu, Ya-Chi, et al.
Published: (2025)
Sketching Low-Rank Plus Diagonal Matrices
by: Fernandez, Andres, et al.
Published: (2025)
by: Fernandez, Andres, et al.
Published: (2025)
Sample-Efficient Hypergradient Estimation for Decentralized Bi-Level Reinforcement Learning
by: Kudo, Mikoto, et al.
Published: (2026)
by: Kudo, Mikoto, et al.
Published: (2026)
Indiscriminate Data Poisoning Attacks on Neural Networks
by: Lu, Yiwei, et al.
Published: (2022)
by: Lu, Yiwei, et al.
Published: (2022)
Non-Stationary Functional Bilevel Optimization
by: Bohne, Jason, et al.
Published: (2026)
by: Bohne, Jason, et al.
Published: (2026)
Functional Bilevel Optimization for Machine Learning
by: Petrulionyte, Ieva, et al.
Published: (2024)
by: Petrulionyte, Ieva, et al.
Published: (2024)
Learning Theory for Kernel Bilevel Optimization
by: Khoury, Fares El, et al.
Published: (2025)
by: Khoury, Fares El, et al.
Published: (2025)
Semiparametric Efficient Bilevel Gradient Estimation
by: Khoury, Fares El, et al.
Published: (2026)
by: Khoury, Fares El, et al.
Published: (2026)
Generalizing the Geometry of Model Merging Through Frechet Averages
by: da Silva, Marvin F., et al.
Published: (2026)
by: da Silva, Marvin F., et al.
Published: (2026)
Bilevel ZOFO: Efficient LLM Fine-Tuning and Meta-Training
by: Shirkavand, Reza, et al.
Published: (2025)
by: Shirkavand, Reza, et al.
Published: (2025)
Natural Hypergradient Descent: Algorithm Design, Convergence Analysis, and Parallel Implementation
by: Kong, Deyi, et al.
Published: (2026)
by: Kong, Deyi, et al.
Published: (2026)
Information-Theoretic Bayesian Optimization for Bilevel Optimization Problems
by: Kanayama, Takuya, et al.
Published: (2025)
by: Kanayama, Takuya, et al.
Published: (2025)
Bayesian Optimization of Bilevel Problems
by: Ekmekcioglu, Omer, et al.
Published: (2024)
by: Ekmekcioglu, Omer, et al.
Published: (2024)
Similar Items
-
Kronecker-factored Approximate Curvature (KFAC) From Scratch
by: Dangel, Felix, et al.
Published: (2025) -
Structured Inverse-Free Natural Gradient: Memory-Efficient & Numerically-Stable KFAC
by: Lin, Wu, et al.
Published: (2023) -
Efficient Curvature-Aware Hypergradient Approximation for Bilevel Optimization
by: Dong, Youran, et al.
Published: (2025) -
Convolutions and More as Einsum: A Tensor Network Perspective with Advances for Second-Order Methods
by: Dangel, Felix
Published: (2023) -
One Sample Fits All: Approximating All Probabilistic Values Simultaneously and Efficiently
by: Li, Weida, et al.
Published: (2024)