Convergence Properties of Stochastic Hypergradients

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Grazzi, Riccardo, Pontil, Massimiliano, Salzo, Saverio
Format: Preprint
Publié: 2020
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910948189536256
author Grazzi, Riccardo
Pontil, Massimiliano
Salzo, Saverio
author_facet Grazzi, Riccardo
Pontil, Massimiliano
Salzo, Saverio
contents Bilevel optimization problems are receiving increasing attention in machine learning as they provide a natural framework for hyperparameter optimization and meta-learning. A key step to tackle these problems is the efficient computation of the gradient of the upper-level objective (hypergradient). In this work, we study stochastic approximation schemes for the hypergradient, which are important when the lower-level problem is empirical risk minimization on a large dataset. The method that we propose is a stochastic variant of the approximate implicit differentiation approach in (Pedregosa, 2016). We provide bounds for the mean square error of the hypergradient approximation, under the assumption that the lower-level problem is accessible only through a stochastic mapping which is a contraction in expectation. In particular, our main bound is agnostic to the choice of the two stochastic solvers employed by the procedure. We provide numerical experiments to support our theoretical analysis and to show the advantage of using stochastic hypergradients in practice.
format Preprint
id arxiv_https___arxiv_org_abs_2011_07122
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle Convergence Properties of Stochastic Hypergradients
Grazzi, Riccardo
Pontil, Massimiliano
Salzo, Saverio
Machine Learning
Bilevel optimization problems are receiving increasing attention in machine learning as they provide a natural framework for hyperparameter optimization and meta-learning. A key step to tackle these problems is the efficient computation of the gradient of the upper-level objective (hypergradient). In this work, we study stochastic approximation schemes for the hypergradient, which are important when the lower-level problem is empirical risk minimization on a large dataset. The method that we propose is a stochastic variant of the approximate implicit differentiation approach in (Pedregosa, 2016). We provide bounds for the mean square error of the hypergradient approximation, under the assumption that the lower-level problem is accessible only through a stochastic mapping which is a contraction in expectation. In particular, our main bound is agnostic to the choice of the two stochastic solvers employed by the procedure. We provide numerical experiments to support our theoretical analysis and to show the advantage of using stochastic hypergradients in practice.
title Convergence Properties of Stochastic Hypergradients
topic Machine Learning
url https://arxiv.org/abs/2011.07122