Outlier-robust Mean Estimation near the Breakdown Point via Sum-of-Squares

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Hongjie, Sridharan, Deepak Narayanan, Steurer, David
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916491006312448
author Chen, Hongjie
Sridharan, Deepak Narayanan
Steurer, David
author_facet Chen, Hongjie
Sridharan, Deepak Narayanan
Steurer, David
contents We revisit the problem of estimating the mean of a high-dimensional distribution in the presence of an $\varepsilon$-fraction of adversarial outliers. When $\varepsilon$ is at most some sufficiently small constant, previous works can achieve optimal error rate efficiently \cite{diakonikolas2018robustly, kothari2018robust}. As $\varepsilon$ approaches the breakdown point $\frac{1}{2}$, all previous algorithms incur either sub-optimal error rates or exponential running time. In this paper we give a new analysis of the canonical sum-of-squares program introduced in \cite{kothari2018robust} and show that this program efficiently achieves optimal error rate for all $\varepsilon \in[0,\frac{1}{2})$. The key ingredient for our results is a new identifiability proof for robust mean estimation that focuses on the overlap between the distributions instead of their statistical distance as in previous works. We capture this proof within the sum-of-squares proof system, thus obtaining efficient algorithms using the sum-of-squares proofs to algorithms paradigm \cite{raghavendra2018high}.
format Preprint
id arxiv_https___arxiv_org_abs_2411_14305
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Outlier-robust Mean Estimation near the Breakdown Point via Sum-of-Squares
Chen, Hongjie
Sridharan, Deepak Narayanan
Steurer, David
Data Structures and Algorithms
Machine Learning
We revisit the problem of estimating the mean of a high-dimensional distribution in the presence of an $\varepsilon$-fraction of adversarial outliers. When $\varepsilon$ is at most some sufficiently small constant, previous works can achieve optimal error rate efficiently \cite{diakonikolas2018robustly, kothari2018robust}. As $\varepsilon$ approaches the breakdown point $\frac{1}{2}$, all previous algorithms incur either sub-optimal error rates or exponential running time. In this paper we give a new analysis of the canonical sum-of-squares program introduced in \cite{kothari2018robust} and show that this program efficiently achieves optimal error rate for all $\varepsilon \in[0,\frac{1}{2})$. The key ingredient for our results is a new identifiability proof for robust mean estimation that focuses on the overlap between the distributions instead of their statistical distance as in previous works. We capture this proof within the sum-of-squares proof system, thus obtaining efficient algorithms using the sum-of-squares proofs to algorithms paradigm \cite{raghavendra2018high}.
title Outlier-robust Mean Estimation near the Breakdown Point via Sum-of-Squares
topic Data Structures and Algorithms
Machine Learning
url https://arxiv.org/abs/2411.14305