A meta-analysis on the performance of machine-learning based language models for sentiment analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rohde, Elena, Klingwort, Jonas, Borgs, Christian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914033061330944
author Rohde, Elena
Klingwort, Jonas
Borgs, Christian
author_facet Rohde, Elena
Klingwort, Jonas
Borgs, Christian
contents This paper presents a meta-analysis evaluating ML performance in sentiment analysis for Twitter data. The study aims to estimate the average performance, assess heterogeneity between and within studies, and analyze how study characteristics influence model performance. Using PRISMA guidelines, we searched academic databases and selected 195 trials from 20 studies with 12 study features. Overall accuracy, the most reported performance metric, was analyzed using double arcsine transformation and a three-level random effects model. The average overall accuracy of the AIC-optimized model was 0.80 [0.76, 0.84]. This paper provides two key insights: 1) Overall accuracy is widely used but often misleading due to its sensitivity to class imbalance and the number of sentiment classes, highlighting the need for normalization. 2) Standardized reporting of model performance, including reporting confusion matrices for independent test sets, is essential for reliable comparisons of ML classifiers across studies, which seems far from common practice.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09728
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A meta-analysis on the performance of machine-learning based language models for sentiment analysis
Rohde, Elena
Klingwort, Jonas
Borgs, Christian
Computation and Language
Machine Learning
Applications
This paper presents a meta-analysis evaluating ML performance in sentiment analysis for Twitter data. The study aims to estimate the average performance, assess heterogeneity between and within studies, and analyze how study characteristics influence model performance. Using PRISMA guidelines, we searched academic databases and selected 195 trials from 20 studies with 12 study features. Overall accuracy, the most reported performance metric, was analyzed using double arcsine transformation and a three-level random effects model. The average overall accuracy of the AIC-optimized model was 0.80 [0.76, 0.84]. This paper provides two key insights: 1) Overall accuracy is widely used but often misleading due to its sensitivity to class imbalance and the number of sentiment classes, highlighting the need for normalization. 2) Standardized reporting of model performance, including reporting confusion matrices for independent test sets, is essential for reliable comparisons of ML classifiers across studies, which seems far from common practice.
title A meta-analysis on the performance of machine-learning based language models for sentiment analysis
topic Computation and Language
Machine Learning
Applications
url https://arxiv.org/abs/2509.09728