SiNFluD: Creating and Evaluating Figurative Language Dataset for Sindhi

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ali, Wazir, Noor, Adeeb, Tumrani, Saifullah
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909029333204992
author Ali, Wazir
Noor, Adeeb
Tumrani, Saifullah
author_facet Ali, Wazir
Noor, Adeeb
Tumrani, Saifullah
contents In this article, we introduce SiNFluD, a novel benchmark dataset for Sindhi figurative language classification. We first collect raw text from various blogs, social media platforms, and literary sources, and subsequently prepare the corpus for annotation. Two native annotators label the data using the Doccano text annotation tool, achieving an inter-annotator agreement of 0.81. We then establish baseline results using 5-fold and 10-fold cross-validation. Finally, we evaluate mBERT, XLM-RoBERTa, and XLM-RoBERTa-XL models, along with SetFit for few-shot fine-tuning of sentence transformers. Among these, the pretrained XLM-RoBERTa-XL achieves the best performance.
format Preprint
id arxiv_https___arxiv_org_abs_2605_01323
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SiNFluD: Creating and Evaluating Figurative Language Dataset for Sindhi
Ali, Wazir
Noor, Adeeb
Tumrani, Saifullah
Computation and Language
Artificial Intelligence
In this article, we introduce SiNFluD, a novel benchmark dataset for Sindhi figurative language classification. We first collect raw text from various blogs, social media platforms, and literary sources, and subsequently prepare the corpus for annotation. Two native annotators label the data using the Doccano text annotation tool, achieving an inter-annotator agreement of 0.81. We then establish baseline results using 5-fold and 10-fold cross-validation. Finally, we evaluate mBERT, XLM-RoBERTa, and XLM-RoBERTa-XL models, along with SetFit for few-shot fine-tuning of sentence transformers. Among these, the pretrained XLM-RoBERTa-XL achieves the best performance.
title SiNFluD: Creating and Evaluating Figurative Language Dataset for Sindhi
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2605.01323