Evaluating Bias and Noise Induced by the U.S. Census Bureau's Privacy Protection Methods

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kenny, Christopher T., McCartan, Cory, Kuriwaki, Shiro, Simko, Tyler, Imai, Kosuke
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916231188054016
author Kenny, Christopher T.
McCartan, Cory
Kuriwaki, Shiro
Simko, Tyler
Imai, Kosuke
author_facet Kenny, Christopher T.
McCartan, Cory
Kuriwaki, Shiro
Simko, Tyler
Imai, Kosuke
contents The United States Census Bureau faces a difficult trade-off between the accuracy of Census statistics and the protection of individual information. We conduct the first independent evaluation of bias and noise induced by the Bureau's two main disclosure avoidance systems: the TopDown algorithm employed for the 2020 Census and the swapping algorithm implemented for the three previous Censuses. Our evaluation leverages the Noisy Measure File (NMF) as well as two independent runs of the TopDown algorithm applied to the 2010 decennial Census. We find that the NMF contains too much noise to be directly useful, especially for Hispanic and multiracial populations. TopDown's post-processing dramatically reduces the NMF noise and produces data whose accuracy is similar to that of swapping. While the estimated errors for both TopDown and swapping algorithms are generally no greater than other sources of Census error, they can be relatively substantial for geographies with small total populations.
format Preprint
id arxiv_https___arxiv_org_abs_2306_07521
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Evaluating Bias and Noise Induced by the U.S. Census Bureau's Privacy Protection Methods
Kenny, Christopher T.
McCartan, Cory
Kuriwaki, Shiro
Simko, Tyler
Imai, Kosuke
Computers and Society
Applications
The United States Census Bureau faces a difficult trade-off between the accuracy of Census statistics and the protection of individual information. We conduct the first independent evaluation of bias and noise induced by the Bureau's two main disclosure avoidance systems: the TopDown algorithm employed for the 2020 Census and the swapping algorithm implemented for the three previous Censuses. Our evaluation leverages the Noisy Measure File (NMF) as well as two independent runs of the TopDown algorithm applied to the 2010 decennial Census. We find that the NMF contains too much noise to be directly useful, especially for Hispanic and multiracial populations. TopDown's post-processing dramatically reduces the NMF noise and produces data whose accuracy is similar to that of swapping. While the estimated errors for both TopDown and swapping algorithms are generally no greater than other sources of Census error, they can be relatively substantial for geographies with small total populations.
title Evaluating Bias and Noise Induced by the U.S. Census Bureau's Privacy Protection Methods
topic Computers and Society
Applications
url https://arxiv.org/abs/2306.07521