Efficient Bilevel Optimization with KFAC-Based Hypergradients

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liao, Disen, Dangel, Felix, Yu, Yaoliang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910088096120832
author Liao, Disen
Dangel, Felix
Yu, Yaoliang
author_facet Liao, Disen
Dangel, Felix
Yu, Yaoliang
contents Bilevel optimization (BO) is widely applicable to many machine learning problems. Scaling BO, however, requires repeatedly computing hypergradients, which involves solving inverse Hessian-vector products (IHVPs). In practice, these operations are often approximated using crude surrogates such as one-step gradient unrolling or identity/short Neumann expansions, which discard curvature information. We build on implicit function theorem-based algorithms and propose to incorporate Kronecker-factored approximate curvature (KFAC), yielding curvature-aware hypergradients with a better performance efficiency trade-off than Conjugate Gradient (CG) or Neumann methods and consistently outperforming unrolling. We evaluate this approach across diverse tasks, including meta-learning and AI safety problems. On models up to BERT, we show that curvature information is valuable at scale, and KFAC can provide it with only modest memory and runtime overhead. Our implementation is available at https://github.com/liaodisen/NeuralBo.
format Preprint
id arxiv_https___arxiv_org_abs_2603_29108
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Efficient Bilevel Optimization with KFAC-Based Hypergradients
Liao, Disen
Dangel, Felix
Yu, Yaoliang
Machine Learning
Bilevel optimization (BO) is widely applicable to many machine learning problems. Scaling BO, however, requires repeatedly computing hypergradients, which involves solving inverse Hessian-vector products (IHVPs). In practice, these operations are often approximated using crude surrogates such as one-step gradient unrolling or identity/short Neumann expansions, which discard curvature information. We build on implicit function theorem-based algorithms and propose to incorporate Kronecker-factored approximate curvature (KFAC), yielding curvature-aware hypergradients with a better performance efficiency trade-off than Conjugate Gradient (CG) or Neumann methods and consistently outperforming unrolling. We evaluate this approach across diverse tasks, including meta-learning and AI safety problems. On models up to BERT, we show that curvature information is valuable at scale, and KFAC can provide it with only modest memory and runtime overhead. Our implementation is available at https://github.com/liaodisen/NeuralBo.
title Efficient Bilevel Optimization with KFAC-Based Hypergradients
topic Machine Learning
url https://arxiv.org/abs/2603.29108