When Bayes goes bad: Weakly-regularized covariate adjustment leads to a biased estimate of prevalence

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kuh, Swen, Kennedy, Lauren, Chen, Qixuan, Gelman, Andrew
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918419143589888
author Kuh, Swen
Kennedy, Lauren
Chen, Qixuan
Gelman, Andrew
author_facet Kuh, Swen
Kennedy, Lauren
Chen, Qixuan
Gelman, Andrew
contents When estimating population prevalence from a non-random sample, it is important to adjust for differences between sample and population. However, adjustment for multiple factors requires analysis that can be difficult to understand and validate. In this manuscript, we explore an unexpected downward trend of estimates when covariates are added sequentially to a Bayesian hierarchical model for the estimation of the prevalence of SARS-CoV-2 specific antibodies in an Australian city in late 2020. We compare our data analysis to results from a simulation study to understand four potential contributors to this effect: (i) correction for differences between sample and population, (ii) rare-events bias in logistic regression, (iii) inclusion of the uncertainty of test sensitivity and specificity in a multilevel model, and (iv) increasing model dimensionality. We find that weak prior distributions on the logistic regression coefficients lead to a systematic increase in the amount of partial pooling across adjustment cells-the prior becomes stronger as model dimensionality increases-which in turn feeds through to the estimated assay specificity, which then feeds back to the model and results in lowering the estimated prevalence. Our paper contributes three elements: (i) immediate and longer-term recommendations for using these types of models, (ii) simulation studies to explore the impact of the contributors to this effect, and (iii) a worked example of investigation of unexpected results in a model with multiple adjustment factors.
format Preprint
id arxiv_https___arxiv_org_abs_2603_29134
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle When Bayes goes bad: Weakly-regularized covariate adjustment leads to a biased estimate of prevalence
Kuh, Swen
Kennedy, Lauren
Chen, Qixuan
Gelman, Andrew
Methodology
Applications
When estimating population prevalence from a non-random sample, it is important to adjust for differences between sample and population. However, adjustment for multiple factors requires analysis that can be difficult to understand and validate. In this manuscript, we explore an unexpected downward trend of estimates when covariates are added sequentially to a Bayesian hierarchical model for the estimation of the prevalence of SARS-CoV-2 specific antibodies in an Australian city in late 2020. We compare our data analysis to results from a simulation study to understand four potential contributors to this effect: (i) correction for differences between sample and population, (ii) rare-events bias in logistic regression, (iii) inclusion of the uncertainty of test sensitivity and specificity in a multilevel model, and (iv) increasing model dimensionality. We find that weak prior distributions on the logistic regression coefficients lead to a systematic increase in the amount of partial pooling across adjustment cells-the prior becomes stronger as model dimensionality increases-which in turn feeds through to the estimated assay specificity, which then feeds back to the model and results in lowering the estimated prevalence. Our paper contributes three elements: (i) immediate and longer-term recommendations for using these types of models, (ii) simulation studies to explore the impact of the contributors to this effect, and (iii) a worked example of investigation of unexpected results in a model with multiple adjustment factors.
title When Bayes goes bad: Weakly-regularized covariate adjustment leads to a biased estimate of prevalence
topic Methodology
Applications
url https://arxiv.org/abs/2603.29134