Comparing Methods for Creating a National Random Sample of Twitter Users
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916154518274048 |
|---|---|
| author | Alizadeh, Meysam Zare, Darya Samei, Zeynab Alizadeh, Mohammadamin Kubli, Mael Aliahmadi, Mohammadhadi Ebrahimi, Sarvenaz Gilardi, Fabrizio |
| author_facet | Alizadeh, Meysam Zare, Darya Samei, Zeynab Alizadeh, Mohammadamin Kubli, Mael Aliahmadi, Mohammadhadi Ebrahimi, Sarvenaz Gilardi, Fabrizio |
| contents | Twitter data has been widely used by researchers across various social and computer science disciplines. A common aim when working with Twitter data is the construction of a random sample of users from a given country. However, while several methods have been proposed in the literature, their comparative performance is mostly unexplored. In this paper, we implement four common methods to collect a random sample of Twitter users in the US: 1% Stream, Bounding Box, Location Query, and Language Query. Then, we compare the methods according to their tweet- and user-level metrics as well as their accuracy in estimating US population with and without using inclusion probabilities of various demographics. Our results show that the 1% Stream method performs differently than others in tweet- and user-level metrics, and best for the construction of a population representative sample. We discuss the conditions under which the 1% Stream method may not be suitable and suggest the Bounding Box method as the second-best method to use. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2402_04879 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Comparing Methods for Creating a National Random Sample of Twitter Users Alizadeh, Meysam Zare, Darya Samei, Zeynab Alizadeh, Mohammadamin Kubli, Mael Aliahmadi, Mohammadhadi Ebrahimi, Sarvenaz Gilardi, Fabrizio Social and Information Networks Twitter data has been widely used by researchers across various social and computer science disciplines. A common aim when working with Twitter data is the construction of a random sample of users from a given country. However, while several methods have been proposed in the literature, their comparative performance is mostly unexplored. In this paper, we implement four common methods to collect a random sample of Twitter users in the US: 1% Stream, Bounding Box, Location Query, and Language Query. Then, we compare the methods according to their tweet- and user-level metrics as well as their accuracy in estimating US population with and without using inclusion probabilities of various demographics. Our results show that the 1% Stream method performs differently than others in tweet- and user-level metrics, and best for the construction of a population representative sample. We discuss the conditions under which the 1% Stream method may not be suitable and suggest the Bounding Box method as the second-best method to use. |
| title | Comparing Methods for Creating a National Random Sample of Twitter Users |
| topic | Social and Information Networks |
| url | https://arxiv.org/abs/2402.04879 |