Comparative Evaluation of Embedding Representations for Financial News Sentiment Analysis

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Roy, Joyjit, Singh, Samaresh Kumar
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910114512896000
author Roy, Joyjit
Singh, Samaresh Kumar
author_facet Roy, Joyjit
Singh, Samaresh Kumar
contents Financial sentiment analysis enhances market understanding. However, standard Natural Language Processing (NLP) approaches encounter significant challenges when applied to small datasets. This study presents a comparative evaluation of embedding-based techniques for financial news sentiment classification in resource-constrained environments. Word2Vec, GloVe, and sentence transformer representations are evaluated in combination with gradient boosting on a manually labeled dataset of 349 financial news headlines. Experimental results identify a substantial gap between validation and test performance. Despite strong validation metrics, models underperform relative to trivial baselines. The analysis indicates that pretrained embeddings yield diminishing returns below a critical data sufficiency threshold. Small validation sets contribute to overfitting during model selection. Practical application is illustrated through weekly sentiment aggregation and narrative summarization for market monitoring. Overall, the findings indicate that embedding quality alone cannot address fundamental data scarcity in sentiment classification. Practitioners with limited labeled data should consider alternative strategies, including few-shot learning, data augmentation, or lexicon-enhanced hybrid methods.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13749
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Comparative Evaluation of Embedding Representations for Financial News Sentiment Analysis
Roy, Joyjit
Singh, Samaresh Kumar
Machine Learning
Artificial Intelligence
Computational Engineering, Finance, and Science
Computers and Society
Software Engineering
Financial sentiment analysis enhances market understanding. However, standard Natural Language Processing (NLP) approaches encounter significant challenges when applied to small datasets. This study presents a comparative evaluation of embedding-based techniques for financial news sentiment classification in resource-constrained environments. Word2Vec, GloVe, and sentence transformer representations are evaluated in combination with gradient boosting on a manually labeled dataset of 349 financial news headlines. Experimental results identify a substantial gap between validation and test performance. Despite strong validation metrics, models underperform relative to trivial baselines. The analysis indicates that pretrained embeddings yield diminishing returns below a critical data sufficiency threshold. Small validation sets contribute to overfitting during model selection. Practical application is illustrated through weekly sentiment aggregation and narrative summarization for market monitoring. Overall, the findings indicate that embedding quality alone cannot address fundamental data scarcity in sentiment classification. Practitioners with limited labeled data should consider alternative strategies, including few-shot learning, data augmentation, or lexicon-enhanced hybrid methods.
title Comparative Evaluation of Embedding Representations for Financial News Sentiment Analysis
topic Machine Learning
Artificial Intelligence
Computational Engineering, Finance, and Science
Computers and Society
Software Engineering
url https://arxiv.org/abs/2512.13749