Why Empirical p-Values Are Not Uniform: Reference Samples, Dependence, and PIT Backtesting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Lis, Jakub
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914569912320000
author Lis, Jakub
author_facet Lis, Jakub
contents Probability integral transforms (PITs) and empirical $p$-values are widely used to assess the calibration of predictive distributions. While exact PIT values are uniformly distributed under correct model specification, practical implementations rely on empirical estimates constructed from finite samples. We show that this estimation step fundamentally alters the statistical structure of the problem. In particular, common-sample and rolling-window implementations introduce dependence and variance distortions that invalidate classical one-sample uniformity tests. When empirical percentiles are conditioned on a shared reference sample, the resulting statistics converge towards a two-sample Kolmogorov--Smirnov regime, while rolling windows induce autocorrelation and variance suppression. Our findings indicate that treating empirical percentiles as independent uniform draws can distort statistical inference and that backtesting procedures based on PITs require revised calibration methods accounting for the underlying two-stage sampling structure.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16221
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Why Empirical p-Values Are Not Uniform: Reference Samples, Dependence, and PIT Backtesting
Lis, Jakub
Methodology
Applications
62P05, 62G10, 62G30, 62G20
Probability integral transforms (PITs) and empirical $p$-values are widely used to assess the calibration of predictive distributions. While exact PIT values are uniformly distributed under correct model specification, practical implementations rely on empirical estimates constructed from finite samples. We show that this estimation step fundamentally alters the statistical structure of the problem. In particular, common-sample and rolling-window implementations introduce dependence and variance distortions that invalidate classical one-sample uniformity tests. When empirical percentiles are conditioned on a shared reference sample, the resulting statistics converge towards a two-sample Kolmogorov--Smirnov regime, while rolling windows induce autocorrelation and variance suppression. Our findings indicate that treating empirical percentiles as independent uniform draws can distort statistical inference and that backtesting procedures based on PITs require revised calibration methods accounting for the underlying two-stage sampling structure.
title Why Empirical p-Values Are Not Uniform: Reference Samples, Dependence, and PIT Backtesting
topic Methodology
Applications
62P05, 62G10, 62G30, 62G20
url https://arxiv.org/abs/2605.16221