Code-Switched Language Identification is Harder Than You Think

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Burchell, Laurie, Birch, Alexandra, Thompson, Robert P., Heafield, Kenneth
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910316854509568
author Burchell, Laurie
Birch, Alexandra
Thompson, Robert P.
Heafield, Kenneth
author_facet Burchell, Laurie
Birch, Alexandra
Thompson, Robert P.
Heafield, Kenneth
contents Code switching (CS) is a very common phenomenon in written and spoken communication but one that is handled poorly by many natural language processing applications. Looking to the application of building CS corpora, we explore CS language identification (LID) for corpus building. We make the task more realistic by scaling it to more languages and considering models with simpler architectures for faster inference. We also reformulate the task as a sentence-level multi-label tagging problem to make it more tractable. Having defined the task, we investigate three reasonable models for this task and define metrics which better reflect desired performance. We present empirical evidence that no current approach is adequate and finally provide recommendations for future work in this area.
format Preprint
id arxiv_https___arxiv_org_abs_2402_01505
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Code-Switched Language Identification is Harder Than You Think
Burchell, Laurie
Birch, Alexandra
Thompson, Robert P.
Heafield, Kenneth
Computation and Language
Code switching (CS) is a very common phenomenon in written and spoken communication but one that is handled poorly by many natural language processing applications. Looking to the application of building CS corpora, we explore CS language identification (LID) for corpus building. We make the task more realistic by scaling it to more languages and considering models with simpler architectures for faster inference. We also reformulate the task as a sentence-level multi-label tagging problem to make it more tractable. Having defined the task, we investigate three reasonable models for this task and define metrics which better reflect desired performance. We present empirical evidence that no current approach is adequate and finally provide recommendations for future work in this area.
title Code-Switched Language Identification is Harder Than You Think
topic Computation and Language
url https://arxiv.org/abs/2402.01505