Saved in:
Bibliographic Details
Main Authors: Thudi, Anvith, Maddison, Chris J.
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2406.01477
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909511068942336
author Thudi, Anvith
Maddison, Chris J.
author_facet Thudi, Anvith
Maddison, Chris J.
contents Machine learning models are often required to perform well across several pre-defined settings, such as a set of user groups. Worst-case performance is a common metric to capture this requirement, and is the objective of group distributionally robust optimization (group DRO). Unfortunately, these methods struggle when the loss is non-convex in the parameters, or the model class is non-parametric. Here, we make a classical move to address this: we reparameterize group DRO from parameter space to function space, which results in a number of advantages. First, we show that group DRO over the space of bounded functions admits a minimax theorem. Second, for cross-entropy and mean squared error, we show that the minimax optimal mixture distribution is the solution of a simple convex optimization problem. Thus, provided one is working with a model class of universal function approximators, group DRO can be solved by a convex optimization problem followed by a classical risk minimization problem. We call our method MixMax. In our experi ments, we found that MixMax matched or outperformed the standard group DRO baselines, and in particular, MixMax improved the performance of XGBoost over the only baseline, data balancing, for variations of the ACSIncome and CelebA annotations datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2406_01477
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MixMax: Distributional Robustness in Function Space via Optimal Data Mixtures
Thudi, Anvith
Maddison, Chris J.
Machine Learning
Machine learning models are often required to perform well across several pre-defined settings, such as a set of user groups. Worst-case performance is a common metric to capture this requirement, and is the objective of group distributionally robust optimization (group DRO). Unfortunately, these methods struggle when the loss is non-convex in the parameters, or the model class is non-parametric. Here, we make a classical move to address this: we reparameterize group DRO from parameter space to function space, which results in a number of advantages. First, we show that group DRO over the space of bounded functions admits a minimax theorem. Second, for cross-entropy and mean squared error, we show that the minimax optimal mixture distribution is the solution of a simple convex optimization problem. Thus, provided one is working with a model class of universal function approximators, group DRO can be solved by a convex optimization problem followed by a classical risk minimization problem. We call our method MixMax. In our experi ments, we found that MixMax matched or outperformed the standard group DRO baselines, and in particular, MixMax improved the performance of XGBoost over the only baseline, data balancing, for variations of the ACSIncome and CelebA annotations datasets.
title MixMax: Distributional Robustness in Function Space via Optimal Data Mixtures
topic Machine Learning
url https://arxiv.org/abs/2406.01477