GRDD+: An Extended Greek Dialectal Dataset with Cross-Architecture Fine-tuning Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chatzikyriakidis, Stergios, Papadakis, Dimitris, Papaioannou, Sevasti-Ioanna, Psaltaki, Erofili
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910035091652608
author Chatzikyriakidis, Stergios
Papadakis, Dimitris
Papaioannou, Sevasti-Ioanna
Psaltaki, Erofili
author_facet Chatzikyriakidis, Stergios
Papadakis, Dimitris
Papaioannou, Sevasti-Ioanna
Psaltaki, Erofili
contents We present an extended Greek Dialectal Dataset (GRDD+) 1that complements the existing GRDD dataset with more data from Cretan, Cypriot, Pontic and Northern Greek, while we add six new varieties: Greco-Corsican, Griko (Southern Italian Greek), Maniot, Heptanesian, Tsakonian, and Katharevusa Greek. The result is a dataset with total size 6,374,939 words and 10 varieties. This is the first dataset with such variation and size to date. We conduct a number of fine-tuning experiments to see the effect of good quality dialectal data on a number of LLMs. We fine-tune three model architectures (Llama-3-8B, Llama-3.1-8B, Krikri-8B) and compare the results to frontier models (Claude-3.7-Sonnet, Gemini-2.5, ChatGPT-5).
format Preprint
id arxiv_https___arxiv_org_abs_2511_03772
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GRDD+: An Extended Greek Dialectal Dataset with Cross-Architecture Fine-tuning Evaluation
Chatzikyriakidis, Stergios
Papadakis, Dimitris
Papaioannou, Sevasti-Ioanna
Psaltaki, Erofili
Computation and Language
We present an extended Greek Dialectal Dataset (GRDD+) 1that complements the existing GRDD dataset with more data from Cretan, Cypriot, Pontic and Northern Greek, while we add six new varieties: Greco-Corsican, Griko (Southern Italian Greek), Maniot, Heptanesian, Tsakonian, and Katharevusa Greek. The result is a dataset with total size 6,374,939 words and 10 varieties. This is the first dataset with such variation and size to date. We conduct a number of fine-tuning experiments to see the effect of good quality dialectal data on a number of LLMs. We fine-tune three model architectures (Llama-3-8B, Llama-3.1-8B, Krikri-8B) and compare the results to frontier models (Claude-3.7-Sonnet, Gemini-2.5, ChatGPT-5).
title GRDD+: An Extended Greek Dialectal Dataset with Cross-Architecture Fine-tuning Evaluation
topic Computation and Language
url https://arxiv.org/abs/2511.03772