GRDD: A Dataset for Greek Dialectal NLP

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chatzikyriakidis, Stergios, Qwaider, Chatrine, Kolokousis, Ilias, Koula, Christina, Papadakis, Dimitris, Sakellariou, Efthymia
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866918160340353024
author Chatzikyriakidis, Stergios
Qwaider, Chatrine
Kolokousis, Ilias
Koula, Christina
Papadakis, Dimitris
Sakellariou, Efthymia
author_facet Chatzikyriakidis, Stergios
Qwaider, Chatrine
Kolokousis, Ilias
Koula, Christina
Papadakis, Dimitris
Sakellariou, Efthymia
contents In this paper, we present a dataset for the computational study of a number of Modern Greek dialects. It consists of raw text data from four dialects of Modern Greek, Cretan, Pontic, Northern Greek and Cypriot Greek. The dataset is of considerable size, albeit imbalanced, and presents the first attempt to create large scale dialectal resources of this type for Modern Greek dialects. We then use the dataset to perform dialect idefntification. We experiment with traditional ML algorithms, as well as simple DL architectures. The results show very good performance on the task, potentially revealing that the dialects in question have distinct enough characteristics allowing even simple ML models to perform well on the task. Error analysis is performed for the top performing algorithms showing that in a number of cases the errors are due to insufficient dataset cleaning.
format Preprint
id arxiv_https___arxiv_org_abs_2308_00802
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle GRDD: A Dataset for Greek Dialectal NLP
Chatzikyriakidis, Stergios
Qwaider, Chatrine
Kolokousis, Ilias
Koula, Christina
Papadakis, Dimitris
Sakellariou, Efthymia
Computation and Language
In this paper, we present a dataset for the computational study of a number of Modern Greek dialects. It consists of raw text data from four dialects of Modern Greek, Cretan, Pontic, Northern Greek and Cypriot Greek. The dataset is of considerable size, albeit imbalanced, and presents the first attempt to create large scale dialectal resources of this type for Modern Greek dialects. We then use the dataset to perform dialect idefntification. We experiment with traditional ML algorithms, as well as simple DL architectures. The results show very good performance on the task, potentially revealing that the dialects in question have distinct enough characteristics allowing even simple ML models to perform well on the task. Error analysis is performed for the top performing algorithms showing that in a number of cases the errors are due to insufficient dataset cleaning.
title GRDD: A Dataset for Greek Dialectal NLP
topic Computation and Language
url https://arxiv.org/abs/2308.00802