Impoverished Language Technology: The Lack of (Social) Class in NLP

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Curry, Amanda Cercas, Talat, Zeerak, Hovy, Dirk
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929266383388672
author Curry, Amanda Cercas
Talat, Zeerak
Hovy, Dirk
author_facet Curry, Amanda Cercas
Talat, Zeerak
Hovy, Dirk
contents Since Labov's (1964) foundational work on the social stratification of language, linguistics has dedicated concerted efforts towards understanding the relationships between socio-demographic factors and language production and perception. Despite the large body of evidence identifying significant relationships between socio-demographic factors and language production, relatively few of these factors have been investigated in the context of NLP technology. While age and gender are well covered, Labov's initial target, socio-economic class, is largely absent. We survey the existing Natural Language Processing (NLP) literature and find that only 20 papers even mention socio-economic status. However, the majority of those papers do not engage with class beyond collecting information of annotator-demographics. Given this research lacuna, we provide a definition of class that can be operationalised by NLP researchers, and argue for including socio-economic class in future language technologies.
format Preprint
id arxiv_https___arxiv_org_abs_2403_03874
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Impoverished Language Technology: The Lack of (Social) Class in NLP
Curry, Amanda Cercas
Talat, Zeerak
Hovy, Dirk
Computation and Language
Artificial Intelligence
Computers and Society
Since Labov's (1964) foundational work on the social stratification of language, linguistics has dedicated concerted efforts towards understanding the relationships between socio-demographic factors and language production and perception. Despite the large body of evidence identifying significant relationships between socio-demographic factors and language production, relatively few of these factors have been investigated in the context of NLP technology. While age and gender are well covered, Labov's initial target, socio-economic class, is largely absent. We survey the existing Natural Language Processing (NLP) literature and find that only 20 papers even mention socio-economic status. However, the majority of those papers do not engage with class beyond collecting information of annotator-demographics. Given this research lacuna, we provide a definition of class that can be operationalised by NLP researchers, and argue for including socio-economic class in future language technologies.
title Impoverished Language Technology: The Lack of (Social) Class in NLP
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2403.03874