Saved in:
Bibliographic Details
Main Authors: Chen, Yiling, Feng, Shi, Kattuman, Paul, Yu, Fang-Yi
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.17085
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909857128382464
author Chen, Yiling
Feng, Shi
Kattuman, Paul
Yu, Fang-Yi
author_facet Chen, Yiling
Feng, Shi
Kattuman, Paul
Yu, Fang-Yi
contents How can we assess the reliability of a dataset without access to ground truth? We introduce the problem of reliability scoring for datasets collected from potentially strategic sources. The true data are unobserved, but we see outcomes of an unknown statistical experiment that depends on them. To benchmark reliability, we define ground-truth-based orderings that capture how much reported data deviate from the truth. We then propose the Gram determinant score, which measures the volume spanned by vectors describing the empirical distribution of the observed data and experiment outcomes. We show that this score preserves several ground-truth based reliability orderings and, uniquely up to scaling, yields the same reliability ranking of datasets regardless of the experiment -- a property we term experiment agnosticism. Experiments on synthetic noise models, CIFAR-10 embeddings, and real employment data demonstrate that the Gram determinant score effectively captures data quality across diverse observation processes.
format Preprint
id arxiv_https___arxiv_org_abs_2510_17085
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data Reliability Scoring
Chen, Yiling
Feng, Shi
Kattuman, Paul
Yu, Fang-Yi
Machine Learning
Computer Science and Game Theory
How can we assess the reliability of a dataset without access to ground truth? We introduce the problem of reliability scoring for datasets collected from potentially strategic sources. The true data are unobserved, but we see outcomes of an unknown statistical experiment that depends on them. To benchmark reliability, we define ground-truth-based orderings that capture how much reported data deviate from the truth. We then propose the Gram determinant score, which measures the volume spanned by vectors describing the empirical distribution of the observed data and experiment outcomes. We show that this score preserves several ground-truth based reliability orderings and, uniquely up to scaling, yields the same reliability ranking of datasets regardless of the experiment -- a property we term experiment agnosticism. Experiments on synthetic noise models, CIFAR-10 embeddings, and real employment data demonstrate that the Gram determinant score effectively captures data quality across diverse observation processes.
title Data Reliability Scoring
topic Machine Learning
Computer Science and Game Theory
url https://arxiv.org/abs/2510.17085