Enhancing Hypergradients Estimation: A Study of Preconditioning and Reparameterization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Zhenzhang, Peyré, Gabriel, Cremers, Daniel, Ablin, Pierre
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916138878763008
author Ye, Zhenzhang
Peyré, Gabriel
Cremers, Daniel
Ablin, Pierre
author_facet Ye, Zhenzhang
Peyré, Gabriel
Cremers, Daniel
Ablin, Pierre
contents Bilevel optimization aims to optimize an outer objective function that depends on the solution to an inner optimization problem. It is routinely used in Machine Learning, notably for hyperparameter tuning. The conventional method to compute the so-called hypergradient of the outer problem is to use the Implicit Function Theorem (IFT). As a function of the error of the inner problem resolution, we study the error of the IFT method. We analyze two strategies to reduce this error: preconditioning the IFT formula and reparameterizing the inner problem. We give a detailed account of the impact of these two modifications on the error, highlighting the role played by higher-order derivatives of the functionals at stake. Our theoretical findings explain when super efficiency, namely reaching an error on the hypergradient that depends quadratically on the error on the inner problem, is achievable and compare the two approaches when this is impossible. Numerical evaluations on hyperparameter tuning for regression problems substantiate our theoretical findings.
format Preprint
id arxiv_https___arxiv_org_abs_2402_16748
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Hypergradients Estimation: A Study of Preconditioning and Reparameterization
Ye, Zhenzhang
Peyré, Gabriel
Cremers, Daniel
Ablin, Pierre
Machine Learning
Bilevel optimization aims to optimize an outer objective function that depends on the solution to an inner optimization problem. It is routinely used in Machine Learning, notably for hyperparameter tuning. The conventional method to compute the so-called hypergradient of the outer problem is to use the Implicit Function Theorem (IFT). As a function of the error of the inner problem resolution, we study the error of the IFT method. We analyze two strategies to reduce this error: preconditioning the IFT formula and reparameterizing the inner problem. We give a detailed account of the impact of these two modifications on the error, highlighting the role played by higher-order derivatives of the functionals at stake. Our theoretical findings explain when super efficiency, namely reaching an error on the hypergradient that depends quadratically on the error on the inner problem, is achievable and compare the two approaches when this is impossible. Numerical evaluations on hyperparameter tuning for regression problems substantiate our theoretical findings.
title Enhancing Hypergradients Estimation: A Study of Preconditioning and Reparameterization
topic Machine Learning
url https://arxiv.org/abs/2402.16748