A Cross-Validation Study of Turkish Sentiment Analysis Datasets and Tools

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Çakıcı, Şevval, Karaduman, Dilara, Çırlan, Mehmet Akif, Hürriyetoğlu, Ali
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916513340981248
author Çakıcı, Şevval
Karaduman, Dilara
Çırlan, Mehmet Akif
Hürriyetoğlu, Ali
author_facet Çakıcı, Şevval
Karaduman, Dilara
Çırlan, Mehmet Akif
Hürriyetoğlu, Ali
contents In recent years, sentiment analysis has gained increasing significance, prompting researchers to explore datasets in various languages, including Turkish. However, the limited availability of Turkish datasets has led to their multifaceted usage in different studies, yielding diverse outcomes. To overcome this challenge, a rigorous review was conducted of research articles published between 2012 and 2022. 31 studies were listed, and 23 Turkish datasets obtained from publicly available sources and email requests used in these studies were collected. We labeled these 31 studies using a taxonomy. We provide a map of sentiment analysis datasets according to this taxonomy in Turkish over 10 years. Moreover, we run state-of-the-art sentiment analysis tools on these datasets and analyzed performance across popular Turkish sentiment datasets. We observed that the performance of the sentiment analysis tools significantly depends on the characteristics of the target text. Our study fosters a more nuanced understanding of sentiment analysis in the Turkish language.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05964
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Cross-Validation Study of Turkish Sentiment Analysis Datasets and Tools
Çakıcı, Şevval
Karaduman, Dilara
Çırlan, Mehmet Akif
Hürriyetoğlu, Ali
Computation and Language
In recent years, sentiment analysis has gained increasing significance, prompting researchers to explore datasets in various languages, including Turkish. However, the limited availability of Turkish datasets has led to their multifaceted usage in different studies, yielding diverse outcomes. To overcome this challenge, a rigorous review was conducted of research articles published between 2012 and 2022. 31 studies were listed, and 23 Turkish datasets obtained from publicly available sources and email requests used in these studies were collected. We labeled these 31 studies using a taxonomy. We provide a map of sentiment analysis datasets according to this taxonomy in Turkish over 10 years. Moreover, we run state-of-the-art sentiment analysis tools on these datasets and analyzed performance across popular Turkish sentiment datasets. We observed that the performance of the sentiment analysis tools significantly depends on the characteristics of the target text. Our study fosters a more nuanced understanding of sentiment analysis in the Turkish language.
title A Cross-Validation Study of Turkish Sentiment Analysis Datasets and Tools
topic Computation and Language
url https://arxiv.org/abs/2412.05964