Split Conformal Prediction under Data Contamination

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Clarkson, Jase, Xu, Wenkai, Cucuringu, Mihai, Swan, Yvik, Reinert, Gesine
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917108821000192
author Clarkson, Jase
Xu, Wenkai
Cucuringu, Mihai
Swan, Yvik
Reinert, Gesine
author_facet Clarkson, Jase
Xu, Wenkai
Cucuringu, Mihai
Swan, Yvik
Reinert, Gesine
contents Conformal prediction is a non-parametric technique for constructing prediction intervals or sets from arbitrary predictive models under the assumption that the data is exchangeable. It is popular as it comes with theoretical guarantees on the marginal coverage of the prediction sets and the split conformal prediction variant has a very low computational cost compared to model training. We study the robustness of split conformal prediction in a data contamination setting, where we assume a small fraction of the calibration scores are drawn from a different distribution than the bulk. We quantify the impact of the corrupted data on the coverage and efficiency of the constructed sets when evaluated on "clean" test points, and verify our results with numerical experiments. Moreover, we propose an adjustment in the classification setting which we call Contamination Robust Conformal Prediction, and verify the efficacy of our approach using both synthetic and real datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2407_07700
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Split Conformal Prediction under Data Contamination
Clarkson, Jase
Xu, Wenkai
Cucuringu, Mihai
Swan, Yvik
Reinert, Gesine
Machine Learning
Conformal prediction is a non-parametric technique for constructing prediction intervals or sets from arbitrary predictive models under the assumption that the data is exchangeable. It is popular as it comes with theoretical guarantees on the marginal coverage of the prediction sets and the split conformal prediction variant has a very low computational cost compared to model training. We study the robustness of split conformal prediction in a data contamination setting, where we assume a small fraction of the calibration scores are drawn from a different distribution than the bulk. We quantify the impact of the corrupted data on the coverage and efficiency of the constructed sets when evaluated on "clean" test points, and verify our results with numerical experiments. Moreover, we propose an adjustment in the classification setting which we call Contamination Robust Conformal Prediction, and verify the efficacy of our approach using both synthetic and real datasets.
title Split Conformal Prediction under Data Contamination
topic Machine Learning
url https://arxiv.org/abs/2407.07700