Complexity of Vector-valued Prediction: From Linear Models to Stochastic Convex Optimization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Schliserman, Matan, Koren, Tomer
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909417541206016
author Schliserman, Matan
Koren, Tomer
author_facet Schliserman, Matan
Koren, Tomer
contents We study the problem of learning vector-valued linear predictors: these are prediction rules parameterized by a matrix that maps an $m$-dimensional feature vector to a $k$-dimensional target. We focus on the fundamental case with a convex and Lipschitz loss function, and show several new theoretical results that shed light on the complexity of this problem and its connection to related learning models. First, we give a tight characterization of the sample complexity of Empirical Risk Minimization (ERM) in this setting, establishing that $\smash{\widetildeΩ}(k/ε^2)$ examples are necessary for ERM to reach $ε$ excess (population) risk; this provides for an exponential improvement over recent results by Magen and Shamir (2023) in terms of the dependence on the target dimension $k$, and matches a classical upper bound due to Maurer (2016). Second, we present a black-box conversion from general $d$-dimensional Stochastic Convex Optimization (SCO) to vector-valued linear prediction, showing that any SCO problem can be embedded as a prediction problem with $k=Θ(d)$ outputs. These results portray the setting of vector-valued linear prediction as bridging between two extensively studied yet disparate learning models: linear models (corresponds to $k=1$) and general $d$-dimensional SCO (with $k=Θ(d)$).
format Preprint
id arxiv_https___arxiv_org_abs_2412_04274
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Complexity of Vector-valued Prediction: From Linear Models to Stochastic Convex Optimization
Schliserman, Matan
Koren, Tomer
Machine Learning
We study the problem of learning vector-valued linear predictors: these are prediction rules parameterized by a matrix that maps an $m$-dimensional feature vector to a $k$-dimensional target. We focus on the fundamental case with a convex and Lipschitz loss function, and show several new theoretical results that shed light on the complexity of this problem and its connection to related learning models. First, we give a tight characterization of the sample complexity of Empirical Risk Minimization (ERM) in this setting, establishing that $\smash{\widetildeΩ}(k/ε^2)$ examples are necessary for ERM to reach $ε$ excess (population) risk; this provides for an exponential improvement over recent results by Magen and Shamir (2023) in terms of the dependence on the target dimension $k$, and matches a classical upper bound due to Maurer (2016). Second, we present a black-box conversion from general $d$-dimensional Stochastic Convex Optimization (SCO) to vector-valued linear prediction, showing that any SCO problem can be embedded as a prediction problem with $k=Θ(d)$ outputs. These results portray the setting of vector-valued linear prediction as bridging between two extensively studied yet disparate learning models: linear models (corresponds to $k=1$) and general $d$-dimensional SCO (with $k=Θ(d)$).
title Complexity of Vector-valued Prediction: From Linear Models to Stochastic Convex Optimization
topic Machine Learning
url https://arxiv.org/abs/2412.04274