Do We Really Even Need Data? A Modern Look at Drawing Inference with Predicted Data

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Salerno, Stephen, Hoffman, Kentaro, Afiaz, Awan, Neufeld, Anna, McCormick, Tyler H., Leek, Jeffrey T.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915656271659008
author Salerno, Stephen
Hoffman, Kentaro
Afiaz, Awan
Neufeld, Anna
McCormick, Tyler H.
Leek, Jeffrey T.
author_facet Salerno, Stephen
Hoffman, Kentaro
Afiaz, Awan
Neufeld, Anna
McCormick, Tyler H.
Leek, Jeffrey T.
contents As artificial intelligence and machine learning tools become more accessible, and scientists face new obstacles to data collection (e.g., rising costs, declining survey response rates), researchers increasingly use predictions from pre-trained algorithms as substitutes for missing or unobserved data. Though appealing for financial and logistical reasons, using standard tools for inference can misrepresent the association between independent variables and the outcome of interest when the true, unobserved outcome is replaced by a predicted value. In this paper, we characterize the statistical challenges inherent to drawing inference with predicted data (IPD) and show that high predictive accuracy does not guarantee valid downstream inference. We show that all such failures reduce to statistical notions of (i) bias, when predictions systematically shift the estimand or distort relationships among variables, and (ii) variance, when uncertainty from the prediction model and the intrinsic variability of the true data are ignored. We then review recent methods for conducting IPD and discuss how this framework is deeply rooted in classical statistical theory. We then comment on some open questions and interesting avenues for future work in this area, and end with some comments on how to use predicted data in scientific studies that is both transparent and statistically principled.
format Preprint
id arxiv_https___arxiv_org_abs_2512_05456
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Do We Really Even Need Data? A Modern Look at Drawing Inference with Predicted Data
Salerno, Stephen
Hoffman, Kentaro
Afiaz, Awan
Neufeld, Anna
McCormick, Tyler H.
Leek, Jeffrey T.
Machine Learning
As artificial intelligence and machine learning tools become more accessible, and scientists face new obstacles to data collection (e.g., rising costs, declining survey response rates), researchers increasingly use predictions from pre-trained algorithms as substitutes for missing or unobserved data. Though appealing for financial and logistical reasons, using standard tools for inference can misrepresent the association between independent variables and the outcome of interest when the true, unobserved outcome is replaced by a predicted value. In this paper, we characterize the statistical challenges inherent to drawing inference with predicted data (IPD) and show that high predictive accuracy does not guarantee valid downstream inference. We show that all such failures reduce to statistical notions of (i) bias, when predictions systematically shift the estimand or distort relationships among variables, and (ii) variance, when uncertainty from the prediction model and the intrinsic variability of the true data are ignored. We then review recent methods for conducting IPD and discuss how this framework is deeply rooted in classical statistical theory. We then comment on some open questions and interesting avenues for future work in this area, and end with some comments on how to use predicted data in scientific studies that is both transparent and statistically principled.
title Do We Really Even Need Data? A Modern Look at Drawing Inference with Predicted Data
topic Machine Learning
url https://arxiv.org/abs/2512.05456