Saved in:
Bibliographic Details
Main Authors: Waldetoft, Hannes, Torgander, Jakob, Magnusson, Måns
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.04643
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • Estimating population parameters in finite populations of text documents can be challenging when obtaining the labels for the target variable requires manual annotation. To address this problem, we combine predictions from a transformer encoder neural network with well-established survey sampling estimators using the model predictions as an auxiliary variable. The applicability is demonstrated in Swedish hate crime statistics based on Swedish police reports. Estimates of the yearly number of hate crimes and the police's under-reporting are derived using the Hansen-Hurwitz estimator, difference estimation, and stratified random sampling estimation. We conclude that if labeled training data is available, the proposed method can provide very efficient estimates with reduced time spent on manual annotation.