From Statistical Features to Contextual Embeddings: A Multi-Paradigm Evaluation of Emotion Classification Models

Fuente: Zenodo
Gespeichert in:
Bibliographische Detailangaben
1. Verfasser: Gregorius, Reynaldi Pratama
Format: Recurso digital
Sprache:Englisch
Veröffentlicht: Zenodo 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866901957494439936
author Gregorius, Reynaldi Pratama
author_facet Gregorius, Reynaldi Pratama
contents <p>This study evaluates multiple model variants across all three paradigms: classical machine learning with statistical feature extraction (Naive Bayes, Logistic Regression, LinearSVC, XGBoost, MLP), deep learning with pre-trained word embeddings (Bidirectional GRU and LSTM with GloVe and FastText), and fine-tuned transformer language models (BERT, RoBERTa, DistilBERT). All models are evaluated on a six-class emotion dataset of 20,000 samples under identical experimental conditions. Results show that the transition from statistical features to GloVe-based embeddings produces the largest single performance gain (+3.6 F1 points), while full transformer fine-tuning yields only a marginal additional improvement (+0.8 points) at substantially higher computational cost. The study provides practical model selection guidance for teams working under real-world resource constraints.</p>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_19028571
institution Zenodo
language eng
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle From Statistical Features to Contextual Embeddings: A Multi-Paradigm Evaluation of Emotion Classification Models
Gregorius, Reynaldi Pratama
emotion classification
natural language processing
transformer fine-tuning
word embeddings
text classification
<p>This study evaluates multiple model variants across all three paradigms: classical machine learning with statistical feature extraction (Naive Bayes, Logistic Regression, LinearSVC, XGBoost, MLP), deep learning with pre-trained word embeddings (Bidirectional GRU and LSTM with GloVe and FastText), and fine-tuned transformer language models (BERT, RoBERTa, DistilBERT). All models are evaluated on a six-class emotion dataset of 20,000 samples under identical experimental conditions. Results show that the transition from statistical features to GloVe-based embeddings produces the largest single performance gain (+3.6 F1 points), while full transformer fine-tuning yields only a marginal additional improvement (+0.8 points) at substantially higher computational cost. The study provides practical model selection guidance for teams working under real-world resource constraints.</p>
title From Statistical Features to Contextual Embeddings: A Multi-Paradigm Evaluation of Emotion Classification Models
topic emotion classification
natural language processing
transformer fine-tuning
word embeddings
text classification
url https://doi.org/10.5281/zenodo.19028571