EuskañolDS: A Naturally Sourced Corpus for Basque-Spanish Code-Switching

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Heredia, Maite, Barnes, Jeremy, Soroa, Aitor
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909478406848512
author Heredia, Maite
Barnes, Jeremy
Soroa, Aitor
author_facet Heredia, Maite
Barnes, Jeremy
Soroa, Aitor
contents Code-switching (CS) remains a significant challenge in Natural Language Processing (NLP), mainly due a lack of relevant data. In the context of the contact between the Basque and Spanish languages in the north of the Iberian Peninsula, CS frequently occurs in both formal and informal spontaneous interactions. However, resources to analyse this phenomenon and support the development and evaluation of models capable of understanding and generating code-switched language for this language pair are almost non-existent. We introduce a first approach to develop a naturally sourced corpus for Basque-Spanish code-switching. Our methodology consists of identifying CS texts from previously available corpora using language identification models, which are then manually validated to obtain a reliable subset of CS instances. We present the properties of our corpus and make it available under the name EuskañolDS.
format Preprint
id arxiv_https___arxiv_org_abs_2502_03188
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EuskañolDS: A Naturally Sourced Corpus for Basque-Spanish Code-Switching
Heredia, Maite
Barnes, Jeremy
Soroa, Aitor
Computation and Language
Artificial Intelligence
Code-switching (CS) remains a significant challenge in Natural Language Processing (NLP), mainly due a lack of relevant data. In the context of the contact between the Basque and Spanish languages in the north of the Iberian Peninsula, CS frequently occurs in both formal and informal spontaneous interactions. However, resources to analyse this phenomenon and support the development and evaluation of models capable of understanding and generating code-switched language for this language pair are almost non-existent. We introduce a first approach to develop a naturally sourced corpus for Basque-Spanish code-switching. Our methodology consists of identifying CS texts from previously available corpora using language identification models, which are then manually validated to obtain a reliable subset of CS instances. We present the properties of our corpus and make it available under the name EuskañolDS.
title EuskañolDS: A Naturally Sourced Corpus for Basque-Spanish Code-Switching
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2502.03188