Comparing Methods for Creating a National Random Sample of Twitter Users

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alizadeh, Meysam, Zare, Darya, Samei, Zeynab, Alizadeh, Mohammadamin, Kubli, Mael, Aliahmadi, Mohammadhadi, Ebrahimi, Sarvenaz, Gilardi, Fabrizio
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916154518274048
author Alizadeh, Meysam
Zare, Darya
Samei, Zeynab
Alizadeh, Mohammadamin
Kubli, Mael
Aliahmadi, Mohammadhadi
Ebrahimi, Sarvenaz
Gilardi, Fabrizio
author_facet Alizadeh, Meysam
Zare, Darya
Samei, Zeynab
Alizadeh, Mohammadamin
Kubli, Mael
Aliahmadi, Mohammadhadi
Ebrahimi, Sarvenaz
Gilardi, Fabrizio
contents Twitter data has been widely used by researchers across various social and computer science disciplines. A common aim when working with Twitter data is the construction of a random sample of users from a given country. However, while several methods have been proposed in the literature, their comparative performance is mostly unexplored. In this paper, we implement four common methods to collect a random sample of Twitter users in the US: 1% Stream, Bounding Box, Location Query, and Language Query. Then, we compare the methods according to their tweet- and user-level metrics as well as their accuracy in estimating US population with and without using inclusion probabilities of various demographics. Our results show that the 1% Stream method performs differently than others in tweet- and user-level metrics, and best for the construction of a population representative sample. We discuss the conditions under which the 1% Stream method may not be suitable and suggest the Bounding Box method as the second-best method to use.
format Preprint
id arxiv_https___arxiv_org_abs_2402_04879
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Comparing Methods for Creating a National Random Sample of Twitter Users
Alizadeh, Meysam
Zare, Darya
Samei, Zeynab
Alizadeh, Mohammadamin
Kubli, Mael
Aliahmadi, Mohammadhadi
Ebrahimi, Sarvenaz
Gilardi, Fabrizio
Social and Information Networks
Twitter data has been widely used by researchers across various social and computer science disciplines. A common aim when working with Twitter data is the construction of a random sample of users from a given country. However, while several methods have been proposed in the literature, their comparative performance is mostly unexplored. In this paper, we implement four common methods to collect a random sample of Twitter users in the US: 1% Stream, Bounding Box, Location Query, and Language Query. Then, we compare the methods according to their tweet- and user-level metrics as well as their accuracy in estimating US population with and without using inclusion probabilities of various demographics. Our results show that the 1% Stream method performs differently than others in tweet- and user-level metrics, and best for the construction of a population representative sample. We discuss the conditions under which the 1% Stream method may not be suitable and suggest the Bounding Box method as the second-best method to use.
title Comparing Methods for Creating a National Random Sample of Twitter Users
topic Social and Information Networks
url https://arxiv.org/abs/2402.04879