Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Vileikytė, Brigita, Lukoševičius, Mantas, Stankevičius, Lukas
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910545621286912
author Vileikytė, Brigita
Lukoševičius, Mantas
Stankevičius, Lukas
author_facet Vileikytė, Brigita
Lukoševičius, Mantas
Stankevičius, Lukas
contents Sentiment analysis is a widely researched area within Natural Language Processing (NLP), attracting significant interest due to the advent of automated solutions. Despite this, the task remains challenging because of the inherent complexity of languages and the subjective nature of sentiments. It is even more challenging for less-studied and less-resourced languages such as Lithuanian. Our review of existing Lithuanian NLP research reveals that traditional machine learning methods and classification algorithms have limited effectiveness for the task. In this work, we address sentiment analysis of Lithuanian five-star-based online reviews from multiple domains that we collect and clean. We apply transformer models to this task for the first time, exploring the capabilities of pre-trained multilingual Large Language Models (LLMs), specifically focusing on fine-tuning BERT and T5 models. Given the inherent difficulty of the task, the fine-tuned models perform quite well, especially when the sentiments themselves are less ambiguous: 80.74% and 89.61% testing recognition accuracy of the most popular one- and five-star reviews respectively. They significantly outperform current commercial state-of-the-art general-purpose LLM GPT-4. We openly share our fine-tuned LLMs online.
format Preprint
id arxiv_https___arxiv_org_abs_2407_19914
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models
Vileikytė, Brigita
Lukoševičius, Mantas
Stankevičius, Lukas
Computation and Language
Information Retrieval
Machine Learning
68T07, 68T50, 68T05,
I.2.6; I.2.7
Sentiment analysis is a widely researched area within Natural Language Processing (NLP), attracting significant interest due to the advent of automated solutions. Despite this, the task remains challenging because of the inherent complexity of languages and the subjective nature of sentiments. It is even more challenging for less-studied and less-resourced languages such as Lithuanian. Our review of existing Lithuanian NLP research reveals that traditional machine learning methods and classification algorithms have limited effectiveness for the task. In this work, we address sentiment analysis of Lithuanian five-star-based online reviews from multiple domains that we collect and clean. We apply transformer models to this task for the first time, exploring the capabilities of pre-trained multilingual Large Language Models (LLMs), specifically focusing on fine-tuning BERT and T5 models. Given the inherent difficulty of the task, the fine-tuned models perform quite well, especially when the sentiments themselves are less ambiguous: 80.74% and 89.61% testing recognition accuracy of the most popular one- and five-star reviews respectively. They significantly outperform current commercial state-of-the-art general-purpose LLM GPT-4. We openly share our fine-tuned LLMs online.
title Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models
topic Computation and Language
Information Retrieval
Machine Learning
68T07, 68T50, 68T05,
I.2.6; I.2.7
url https://arxiv.org/abs/2407.19914