Exploring Data and Parameter Efficient Strategies for Arabic Dialect Identifications

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kanjirangat, Vani, Dolamic, Ljiljana, Rinaldi, Fabio
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912591307079680
author Kanjirangat, Vani
Dolamic, Ljiljana
Rinaldi, Fabio
author_facet Kanjirangat, Vani
Dolamic, Ljiljana
Rinaldi, Fabio
contents This paper discusses our exploration of different data-efficient and parameter-efficient approaches to Arabic Dialect Identification (ADI). In particular, we investigate various soft-prompting strategies, including prefix-tuning, prompt-tuning, P-tuning, and P-tuning V2, as well as LoRA reparameterizations. For the data-efficient strategy, we analyze hard prompting with zero-shot and few-shot inferences to analyze the dialect identification capabilities of Large Language Models (LLMs). For the parameter-efficient PEFT approaches, we conducted our experiments using Arabic-specific encoder models on several major datasets. We also analyzed the n-shot inferences on open-source decoder-only models, a general multilingual model (Phi-3.5), and an Arabic-specific one(SILMA). We observed that the LLMs generally struggle to differentiate the dialectal nuances in the few-shot or zero-shot setups. The soft-prompted encoder variants perform better, while the LoRA-based fine-tuned models perform best, even surpassing full fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13775
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Data and Parameter Efficient Strategies for Arabic Dialect Identifications
Kanjirangat, Vani
Dolamic, Ljiljana
Rinaldi, Fabio
Computation and Language
Artificial Intelligence
This paper discusses our exploration of different data-efficient and parameter-efficient approaches to Arabic Dialect Identification (ADI). In particular, we investigate various soft-prompting strategies, including prefix-tuning, prompt-tuning, P-tuning, and P-tuning V2, as well as LoRA reparameterizations. For the data-efficient strategy, we analyze hard prompting with zero-shot and few-shot inferences to analyze the dialect identification capabilities of Large Language Models (LLMs). For the parameter-efficient PEFT approaches, we conducted our experiments using Arabic-specific encoder models on several major datasets. We also analyzed the n-shot inferences on open-source decoder-only models, a general multilingual model (Phi-3.5), and an Arabic-specific one(SILMA). We observed that the LLMs generally struggle to differentiate the dialectal nuances in the few-shot or zero-shot setups. The soft-prompted encoder variants perform better, while the LoRA-based fine-tuned models perform best, even surpassing full fine-tuning.
title Exploring Data and Parameter Efficient Strategies for Arabic Dialect Identifications
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2509.13775