Conformal Data Contamination Tests for Trading or Sharing of Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vejling, Martin V., Pandey, Shashi Raj, Biscio, Christophe A. N., Popovski, Petar
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911063651385344
author Vejling, Martin V.
Pandey, Shashi Raj
Biscio, Christophe A. N.
Popovski, Petar
author_facet Vejling, Martin V.
Pandey, Shashi Raj
Biscio, Christophe A. N.
Popovski, Petar
contents The amount of quality data in many machine learning tasks is limited to what is available locally to data owners. The set of quality data can be expanded through trading or sharing with external data agents. However, data buyers need quality guarantees before purchasing, as external data may be contaminated or irrelevant to their specific learning task. Previous works primarily rely on distributional assumptions about data from different agents, relegating quality checks to post-hoc steps involving costly data valuation procedures. We propose a distribution-free, contamination-aware data-sharing framework that identifies external data agents whose data is most valuable for model personalization. To achieve this, we introduce novel two-sample testing procedures, grounded in rigorous theoretical foundations for conformal outlier detection, to determine whether an agent's data exceeds a contamination threshold. The proposed tests, termed conformal data contamination tests, remain valid under arbitrary contamination levels while enabling false discovery rate control via the Benjamini-Hochberg procedure. Empirical evaluations across diverse collaborative learning scenarios demonstrate the robustness and effectiveness of our approach. Overall, the conformal data contamination test distinguishes itself as a generic procedure for aggregating data with statistically rigorous quality guarantees.
format Preprint
id arxiv_https___arxiv_org_abs_2507_13835
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Conformal Data Contamination Tests for Trading or Sharing of Data
Vejling, Martin V.
Pandey, Shashi Raj
Biscio, Christophe A. N.
Popovski, Petar
Machine Learning
The amount of quality data in many machine learning tasks is limited to what is available locally to data owners. The set of quality data can be expanded through trading or sharing with external data agents. However, data buyers need quality guarantees before purchasing, as external data may be contaminated or irrelevant to their specific learning task. Previous works primarily rely on distributional assumptions about data from different agents, relegating quality checks to post-hoc steps involving costly data valuation procedures. We propose a distribution-free, contamination-aware data-sharing framework that identifies external data agents whose data is most valuable for model personalization. To achieve this, we introduce novel two-sample testing procedures, grounded in rigorous theoretical foundations for conformal outlier detection, to determine whether an agent's data exceeds a contamination threshold. The proposed tests, termed conformal data contamination tests, remain valid under arbitrary contamination levels while enabling false discovery rate control via the Benjamini-Hochberg procedure. Empirical evaluations across diverse collaborative learning scenarios demonstrate the robustness and effectiveness of our approach. Overall, the conformal data contamination test distinguishes itself as a generic procedure for aggregating data with statistically rigorous quality guarantees.
title Conformal Data Contamination Tests for Trading or Sharing of Data
topic Machine Learning
url https://arxiv.org/abs/2507.13835