Conformal Validity Guarantees Exist for Any Data Distribution (and How to Find Them)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Prinster, Drew, Stanton, Samuel, Liu, Anqi, Saria, Suchi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909217558888448
author Prinster, Drew
Stanton, Samuel
Liu, Anqi
Saria, Suchi
author_facet Prinster, Drew
Stanton, Samuel
Liu, Anqi
Saria, Suchi
contents As artificial intelligence (AI) / machine learning (ML) gain widespread adoption, practitioners are increasingly seeking means to quantify and control the risk these systems incur. This challenge is especially salient when such systems have autonomy to collect their own data, such as in black-box optimization and active learning, where their actions induce sequential feedback-loop shifts in the data distribution. Conformal prediction is a promising approach to uncertainty and risk quantification, but prior variants' validity guarantees have assumed some form of ``quasi-exchangeability'' on the data distribution, thereby excluding many types of sequential shifts. In this paper we prove that conformal prediction can theoretically be extended to \textit{any} joint data distribution, not just exchangeable or quasi-exchangeable ones. Although the most general case is exceedingly impractical to compute, for concrete practical applications we outline a procedure for deriving specific conformal algorithms for any data distribution, and we use this procedure to derive tractable algorithms for a series of AI/ML-agent-induced covariate shifts. We evaluate the proposed algorithms empirically on synthetic black-box optimization and active learning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_06627
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Conformal Validity Guarantees Exist for Any Data Distribution (and How to Find Them)
Prinster, Drew
Stanton, Samuel
Liu, Anqi
Saria, Suchi
Machine Learning
Artificial Intelligence
As artificial intelligence (AI) / machine learning (ML) gain widespread adoption, practitioners are increasingly seeking means to quantify and control the risk these systems incur. This challenge is especially salient when such systems have autonomy to collect their own data, such as in black-box optimization and active learning, where their actions induce sequential feedback-loop shifts in the data distribution. Conformal prediction is a promising approach to uncertainty and risk quantification, but prior variants' validity guarantees have assumed some form of ``quasi-exchangeability'' on the data distribution, thereby excluding many types of sequential shifts. In this paper we prove that conformal prediction can theoretically be extended to \textit{any} joint data distribution, not just exchangeable or quasi-exchangeable ones. Although the most general case is exceedingly impractical to compute, for concrete practical applications we outline a procedure for deriving specific conformal algorithms for any data distribution, and we use this procedure to derive tractable algorithms for a series of AI/ML-agent-induced covariate shifts. We evaluate the proposed algorithms empirically on synthetic black-box optimization and active learning tasks.
title Conformal Validity Guarantees Exist for Any Data Distribution (and How to Find Them)
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2405.06627