Data-Driven Strategies for Detecting and Sampling Misrepresented Subgroups

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lancia, G., Mecatti, F., Riccomagno, E.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914246058573824
author Lancia, G.
Mecatti, F.
Riccomagno, E.
author_facet Lancia, G.
Mecatti, F.
Riccomagno, E.
contents Economic policy research frequently examines population well-being, with a particular focus on the relationships between unequal living conditions, low educational attainment, and social exclusion. Sample surveys, such as EU-SILC, are widely used for this purpose and inform public policy; yet, their sampling designs may fail to adequately represent rare, hard-to-sample, or under-covered subgroups. This limitation can hinder socio-demographic analyses and evidence-based policy design. We propose a generalisable approach based on univariate and multivariate unsupervised learning techniques to detect outliers in survey data that may signal under-represented subgroups. Identified groups can then be characterised to inform targeted resampling strategies that improve survey inclusiveness. An empirical application using the 2019 EU-SILC data for the Italian region of Liguria shows that citizenship, material deprivation, large household size, and economic vulnerability are key indicators of under-representation.
format Preprint
id arxiv_https___arxiv_org_abs_2405_01342
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data-Driven Strategies for Detecting and Sampling Misrepresented Subgroups
Lancia, G.
Mecatti, F.
Riccomagno, E.
Applications
Computation
Other Statistics
Economic policy research frequently examines population well-being, with a particular focus on the relationships between unequal living conditions, low educational attainment, and social exclusion. Sample surveys, such as EU-SILC, are widely used for this purpose and inform public policy; yet, their sampling designs may fail to adequately represent rare, hard-to-sample, or under-covered subgroups. This limitation can hinder socio-demographic analyses and evidence-based policy design. We propose a generalisable approach based on univariate and multivariate unsupervised learning techniques to detect outliers in survey data that may signal under-represented subgroups. Identified groups can then be characterised to inform targeted resampling strategies that improve survey inclusiveness. An empirical application using the 2019 EU-SILC data for the Italian region of Liguria shows that citizenship, material deprivation, large household size, and economic vulnerability are key indicators of under-representation.
title Data-Driven Strategies for Detecting and Sampling Misrepresented Subgroups
topic Applications
Computation
Other Statistics
url https://arxiv.org/abs/2405.01342