Automatically detecting codeswitching between Nigerian Pidgin English and English

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ogbonda, Godspower, Teahan, William
Format: Recurso digital
Veröffentlicht: Zenodo 2026
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866902093471678464
author Ogbonda, Godspower
Teahan, William
author_facet Ogbonda, Godspower
Teahan, William
contents <p>This dataset contains manually annotated ground truth files for <br>evaluating code-switching detection between Nigerian Pidgin English <br>(NPE) and Standard English. It was developed as part of a study <br>evaluating traditional machine learning classifiers, a BiLSTM deep <br>learning model, and compression-based PPM models via the Tawa toolkit.</p> <p>Two annotation formats are provided:</p> <p>ml_ground_truth: XML-tagged annotations used to evaluate traditional <br>machine learning and deep learning models, where pidgin and English <br>Segments are marked using XML-style tags.</p> <p>tawa_ground_truth: Character-level annotations used to evaluate the <br>Tawa PPM compression-based models, where P denotes Nigerian Pidgin <br>English and E denotes Standard English.</p> <p>All annotations were manually verified by three native speakers of <br>Nigerian Pidgin English. The ground truth was withheld entirely from model training to ensure unbiased <br>evaluation. Full methodological details are provided in the associated <br>publication.</p> <p>Note: This ground truth dataset is released independently of the full <br>Mixed corpus. The complete corpus cannot be publicly distributed due <br>to copyright restrictions on a subset of the source materials.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_20117988
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle Automatically detecting codeswitching between Nigerian Pidgin English and English
Ogbonda, Godspower
Teahan, William
<p>This dataset contains manually annotated ground truth files for <br>evaluating code-switching detection between Nigerian Pidgin English <br>(NPE) and Standard English. It was developed as part of a study <br>evaluating traditional machine learning classifiers, a BiLSTM deep <br>learning model, and compression-based PPM models via the Tawa toolkit.</p> <p>Two annotation formats are provided:</p> <p>ml_ground_truth: XML-tagged annotations used to evaluate traditional <br>machine learning and deep learning models, where pidgin and English <br>Segments are marked using XML-style tags.</p> <p>tawa_ground_truth: Character-level annotations used to evaluate the <br>Tawa PPM compression-based models, where P denotes Nigerian Pidgin <br>English and E denotes Standard English.</p> <p>All annotations were manually verified by three native speakers of <br>Nigerian Pidgin English. The ground truth was withheld entirely from model training to ensure unbiased <br>evaluation. Full methodological details are provided in the associated <br>publication.</p> <p>Note: This ground truth dataset is released independently of the full <br>Mixed corpus. The complete corpus cannot be publicly distributed due <br>to copyright restrictions on a subset of the source materials.</p>
title Automatically detecting codeswitching between Nigerian Pidgin English and English
url https://doi.org/10.5281/zenodo.20117988