Exploring Subjective Tasks in Farsi: A Survey Analysis and Evaluation of Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rooein, Donya, Plaza-del-Arco, Flor Miriam, Nozza, Debora, Hovy, Dirk
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915483083603968
author Rooein, Donya
Plaza-del-Arco, Flor Miriam
Nozza, Debora
Hovy, Dirk
author_facet Rooein, Donya
Plaza-del-Arco, Flor Miriam
Nozza, Debora
Hovy, Dirk
contents Given Farsi's speaker base of over 127 million people and the growing availability of digital text, including more than 1.3 million articles on Wikipedia, it is considered a middle-resource language. However, this label quickly crumbles when the situation is examined more closely. We focus on three subjective tasks (Sentiment Analysis, Emotion Analysis, and Toxicity Detection) and find significant challenges in data availability and quality, despite the overall increase in data availability. We review 110 publications on subjective tasks in Farsi and observe a lack of publicly available datasets. Furthermore, existing datasets often lack essential demographic factors, such as age and gender, that are crucial for accurately modeling subjectivity in language. When evaluating prediction models using the few available datasets, the results are highly unstable across both datasets and models. Our findings indicate that the volume of data is insufficient to significantly improve a language's prospects in NLP.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05719
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Subjective Tasks in Farsi: A Survey Analysis and Evaluation of Language Models
Rooein, Donya
Plaza-del-Arco, Flor Miriam
Nozza, Debora
Hovy, Dirk
Computation and Language
Given Farsi's speaker base of over 127 million people and the growing availability of digital text, including more than 1.3 million articles on Wikipedia, it is considered a middle-resource language. However, this label quickly crumbles when the situation is examined more closely. We focus on three subjective tasks (Sentiment Analysis, Emotion Analysis, and Toxicity Detection) and find significant challenges in data availability and quality, despite the overall increase in data availability. We review 110 publications on subjective tasks in Farsi and observe a lack of publicly available datasets. Furthermore, existing datasets often lack essential demographic factors, such as age and gender, that are crucial for accurately modeling subjectivity in language. When evaluating prediction models using the few available datasets, the results are highly unstable across both datasets and models. Our findings indicate that the volume of data is insufficient to significantly improve a language's prospects in NLP.
title Exploring Subjective Tasks in Farsi: A Survey Analysis and Evaluation of Language Models
topic Computation and Language
url https://arxiv.org/abs/2509.05719