Splits! Flexible Sociocultural Linguistic Investigation at Scale

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Caplan, Eylon, Chakraborty, Tania, Goldwasser, Dan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911578673119232
author Caplan, Eylon
Chakraborty, Tania
Goldwasser, Dan
author_facet Caplan, Eylon
Chakraborty, Tania
Goldwasser, Dan
contents Variation in language use, shaped by speakers' sociocultural background and specific context of use, offers a rich lens into cultural perspectives, values, and opinions. For example, Chinese students discuss "healthy eating" with words like "timing," "regularity," and "digestion," whereas Americans use vocabulary like "balancing food groups" and "avoiding fat and sugar," reflecting distinct cultural models of nutrition. The computational study of these Sociocultural Linguistic Phenomena (SLP) has traditionally been done in NLP via tailored analyses of specific groups or topics, requiring specialized data collection and experimental operationalization--a process not well-suited to quick hypothesis exploration and prototyping. To address this, we propose constructing a "sandbox" designed for systematic and flexible sociolinguistic research. Using our method, we construct a demographically/topically split Reddit dataset, Splits!, validated by self-identification and by replicating several known SLPs from existing literature. We showcase the sandbox's utility with a scalable, two-stage process that filters large collections of "potential" SLPs (PSLPs) to surface the most promising candidates for deeper, qualitative investigation.
format Preprint
id arxiv_https___arxiv_org_abs_2504_04640
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Splits! Flexible Sociocultural Linguistic Investigation at Scale
Caplan, Eylon
Chakraborty, Tania
Goldwasser, Dan
Computation and Language
Artificial Intelligence
Variation in language use, shaped by speakers' sociocultural background and specific context of use, offers a rich lens into cultural perspectives, values, and opinions. For example, Chinese students discuss "healthy eating" with words like "timing," "regularity," and "digestion," whereas Americans use vocabulary like "balancing food groups" and "avoiding fat and sugar," reflecting distinct cultural models of nutrition. The computational study of these Sociocultural Linguistic Phenomena (SLP) has traditionally been done in NLP via tailored analyses of specific groups or topics, requiring specialized data collection and experimental operationalization--a process not well-suited to quick hypothesis exploration and prototyping. To address this, we propose constructing a "sandbox" designed for systematic and flexible sociolinguistic research. Using our method, we construct a demographically/topically split Reddit dataset, Splits!, validated by self-identification and by replicating several known SLPs from existing literature. We showcase the sandbox's utility with a scalable, two-stage process that filters large collections of "potential" SLPs (PSLPs) to surface the most promising candidates for deeper, qualitative investigation.
title Splits! Flexible Sociocultural Linguistic Investigation at Scale
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.04640