Detecting Linguistic Diversity on Social Media

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wong, Sidney, Adams, Benjamin, Dunn, Jonathan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866929735618002944
author Wong, Sidney
Adams, Benjamin
Dunn, Jonathan
author_facet Wong, Sidney
Adams, Benjamin
Dunn, Jonathan
contents This chapter explores the efficacy of using social media data to examine changing linguistic behaviour of a place. We focus our investigation on Aotearoa New Zealand where official statistics from the census is the only source of language use data. We use published census data as the ground truth and the social media sub-corpus from the Corpus of Global Language Use as our alternative data source. We use place as the common denominator between the two data sources. We identify the language conditions of each tweet in the social media data set and validated our results with two language identification models. We then compare levels of linguistic diversity at national, regional, and local geographies. The results suggest that social media language data has the possibility to provide a rich source of spatial and temporal insights on the linguistic profile of a place. We show that social media is sensitive to demographic and sociopolitical changes within a language and at low-level regional and local geographies.
format Preprint
id arxiv_https___arxiv_org_abs_2502_21224
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Detecting Linguistic Diversity on Social Media
Wong, Sidney
Adams, Benjamin
Dunn, Jonathan
Computation and Language
This chapter explores the efficacy of using social media data to examine changing linguistic behaviour of a place. We focus our investigation on Aotearoa New Zealand where official statistics from the census is the only source of language use data. We use published census data as the ground truth and the social media sub-corpus from the Corpus of Global Language Use as our alternative data source. We use place as the common denominator between the two data sources. We identify the language conditions of each tweet in the social media data set and validated our results with two language identification models. We then compare levels of linguistic diversity at national, regional, and local geographies. The results suggest that social media language data has the possibility to provide a rich source of spatial and temporal insights on the linguistic profile of a place. We show that social media is sensitive to demographic and sociopolitical changes within a language and at low-level regional and local geographies.
title Detecting Linguistic Diversity on Social Media
topic Computation and Language
url https://arxiv.org/abs/2502.21224