PoPreRo: A New Dataset for Popularity Prediction of Romanian Reddit Posts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rogoz, Ana-Cristina, Nechita, Maria Ilinca, Ionescu, Radu Tudor
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910711845748736
author Rogoz, Ana-Cristina
Nechita, Maria Ilinca
Ionescu, Radu Tudor
author_facet Rogoz, Ana-Cristina
Nechita, Maria Ilinca
Ionescu, Radu Tudor
contents We introduce PoPreRo, the first dataset for Popularity Prediction of Romanian posts collected from Reddit. The PoPreRo dataset includes a varied compilation of post samples from five distinct subreddits of Romania, totaling 28,107 data samples. Along with our novel dataset, we introduce a set of competitive models to be used as baselines for future research. Interestingly, the top-scoring model achieves an accuracy of 61.35% and a macro F1 score of 60.60% on the test set, indicating that the popularity prediction task on PoPreRo is very challenging. Further investigations based on few-shot prompting the Falcon-7B Large Language Model also point in the same direction. We thus believe that PoPreRo is a valuable resource that can be used to evaluate models on predicting the popularity of social media posts in Romanian. We release our dataset at https://github.com/ana-rogoz/PoPreRo.
format Preprint
id arxiv_https___arxiv_org_abs_2407_04541
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PoPreRo: A New Dataset for Popularity Prediction of Romanian Reddit Posts
Rogoz, Ana-Cristina
Nechita, Maria Ilinca
Ionescu, Radu Tudor
Computation and Language
Artificial Intelligence
Machine Learning
We introduce PoPreRo, the first dataset for Popularity Prediction of Romanian posts collected from Reddit. The PoPreRo dataset includes a varied compilation of post samples from five distinct subreddits of Romania, totaling 28,107 data samples. Along with our novel dataset, we introduce a set of competitive models to be used as baselines for future research. Interestingly, the top-scoring model achieves an accuracy of 61.35% and a macro F1 score of 60.60% on the test set, indicating that the popularity prediction task on PoPreRo is very challenging. Further investigations based on few-shot prompting the Falcon-7B Large Language Model also point in the same direction. We thus believe that PoPreRo is a valuable resource that can be used to evaluate models on predicting the popularity of social media posts in Romanian. We release our dataset at https://github.com/ana-rogoz/PoPreRo.
title PoPreRo: A New Dataset for Popularity Prediction of Romanian Reddit Posts
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.04541