Distributionally Robust Optimization with Adversarial Data Contamination

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Shuyao, Diakonikolas, Ilias, Diakonikolas, Jelena
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915591138312192
author Li, Shuyao
Diakonikolas, Ilias
Diakonikolas, Jelena
author_facet Li, Shuyao
Diakonikolas, Ilias
Diakonikolas, Jelena
contents Distributionally Robust Optimization (DRO) provides a framework for decision-making under distributional uncertainty, yet its effectiveness can be compromised by outliers in the training data. This paper introduces a principled approach to simultaneously address both challenges. We focus on optimizing Wasserstein-1 DRO objectives for generalized linear models with convex Lipschitz loss functions, where an $ε$-fraction of the training data is adversarially corrupted. Our primary contribution lies in a novel modeling framework that integrates robustness against training data contamination with robustness against distributional shifts, alongside an efficient algorithm inspired by robust statistics to solve the resulting optimization problem. We prove that our method achieves an estimation error of $O(\sqrtε)$ for the true DRO objective value using only the contaminated data under the bounded covariance assumption. This work establishes the first rigorous guarantees, supported by efficient computation, for learning under the dual challenges of data contamination and distributional shifts.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10718
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Distributionally Robust Optimization with Adversarial Data Contamination
Li, Shuyao
Diakonikolas, Ilias
Diakonikolas, Jelena
Machine Learning
Data Structures and Algorithms
Optimization and Control
Distributionally Robust Optimization (DRO) provides a framework for decision-making under distributional uncertainty, yet its effectiveness can be compromised by outliers in the training data. This paper introduces a principled approach to simultaneously address both challenges. We focus on optimizing Wasserstein-1 DRO objectives for generalized linear models with convex Lipschitz loss functions, where an $ε$-fraction of the training data is adversarially corrupted. Our primary contribution lies in a novel modeling framework that integrates robustness against training data contamination with robustness against distributional shifts, alongside an efficient algorithm inspired by robust statistics to solve the resulting optimization problem. We prove that our method achieves an estimation error of $O(\sqrtε)$ for the true DRO objective value using only the contaminated data under the bounded covariance assumption. This work establishes the first rigorous guarantees, supported by efficient computation, for learning under the dual challenges of data contamination and distributional shifts.
title Distributionally Robust Optimization with Adversarial Data Contamination
topic Machine Learning
Data Structures and Algorithms
Optimization and Control
url https://arxiv.org/abs/2507.10718