Comparing Methods for Creating a National Random Sample of Twitter Users

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Alizadeh, Meysam, Zare, Darya, Samei, Zeynab, Alizadeh, Mohammadamin, Kubli, Mael, Aliahmadi, Mohammadhadi, Ebrahimi, Sarvenaz, Gilardi, Fabrizio
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916154518274048
author Alizadeh, Meysam
Zare, Darya
Samei, Zeynab
Alizadeh, Mohammadamin
Kubli, Mael
Aliahmadi, Mohammadhadi
Ebrahimi, Sarvenaz
Gilardi, Fabrizio
author_facet Alizadeh, Meysam
Zare, Darya
Samei, Zeynab
Alizadeh, Mohammadamin
Kubli, Mael
Aliahmadi, Mohammadhadi
Ebrahimi, Sarvenaz
Gilardi, Fabrizio
contents Twitter data has been widely used by researchers across various social and computer science disciplines. A common aim when working with Twitter data is the construction of a random sample of users from a given country. However, while several methods have been proposed in the literature, their comparative performance is mostly unexplored. In this paper, we implement four common methods to collect a random sample of Twitter users in the US: 1% Stream, Bounding Box, Location Query, and Language Query. Then, we compare the methods according to their tweet- and user-level metrics as well as their accuracy in estimating US population with and without using inclusion probabilities of various demographics. Our results show that the 1% Stream method performs differently than others in tweet- and user-level metrics, and best for the construction of a population representative sample. We discuss the conditions under which the 1% Stream method may not be suitable and suggest the Bounding Box method as the second-best method to use.
format Preprint
id arxiv_https___arxiv_org_abs_2402_04879
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Comparing Methods for Creating a National Random Sample of Twitter Users
Alizadeh, Meysam
Zare, Darya
Samei, Zeynab
Alizadeh, Mohammadamin
Kubli, Mael
Aliahmadi, Mohammadhadi
Ebrahimi, Sarvenaz
Gilardi, Fabrizio
Social and Information Networks
Twitter data has been widely used by researchers across various social and computer science disciplines. A common aim when working with Twitter data is the construction of a random sample of users from a given country. However, while several methods have been proposed in the literature, their comparative performance is mostly unexplored. In this paper, we implement four common methods to collect a random sample of Twitter users in the US: 1% Stream, Bounding Box, Location Query, and Language Query. Then, we compare the methods according to their tweet- and user-level metrics as well as their accuracy in estimating US population with and without using inclusion probabilities of various demographics. Our results show that the 1% Stream method performs differently than others in tweet- and user-level metrics, and best for the construction of a population representative sample. We discuss the conditions under which the 1% Stream method may not be suitable and suggest the Bounding Box method as the second-best method to use.
title Comparing Methods for Creating a National Random Sample of Twitter Users
topic Social and Information Networks
url https://arxiv.org/abs/2402.04879