AgoraSpeech: A multi-annotated comprehensive dataset of political discourse through the lens of humans and AI

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sermpezis, Pavlos, Karamanidis, Stelios, Paraschou, Eva, Dimitriadis, Ilias, Yfantidou, Sofia, Kouskouveli, Filitsa-Ioanna, Troboukis, Thanasis, Kiki, Kelly, Galanopoulos, Antonis, Vakali, Athena
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909454346223616
author Sermpezis, Pavlos
Karamanidis, Stelios
Paraschou, Eva
Dimitriadis, Ilias
Yfantidou, Sofia
Kouskouveli, Filitsa-Ioanna
Troboukis, Thanasis
Kiki, Kelly
Galanopoulos, Antonis
Vakali, Athena
author_facet Sermpezis, Pavlos
Karamanidis, Stelios
Paraschou, Eva
Dimitriadis, Ilias
Yfantidou, Sofia
Kouskouveli, Filitsa-Ioanna
Troboukis, Thanasis
Kiki, Kelly
Galanopoulos, Antonis
Vakali, Athena
contents Political discourse datasets are important for gaining political insights, analyzing communication strategies or social science phenomena. Although numerous political discourse corpora exist, comprehensive, high-quality, annotated datasets are scarce. This is largely due to the substantial manual effort, multidisciplinarity, and expertise required for the nuanced annotation of rhetorical strategies and ideological contexts. In this paper, we present AgoraSpeech, a meticulously curated, high-quality dataset of 171 political speeches from six parties during the Greek national elections in 2023. The dataset includes annotations (per paragraph) for six natural language processing (NLP) tasks: text classification, topic identification, sentiment analysis, named entity recognition, polarization and populism detection. A two-step annotation was employed, starting with ChatGPT-generated annotations and followed by exhaustive human-in-the-loop validation. The dataset was initially used in a case study to provide insights during the pre-election period. However, it has general applicability by serving as a rich source of information for political and social scientists, journalists, or data scientists, while it can be used for benchmarking and fine-tuning NLP and large language models (LLMs).
format Preprint
id arxiv_https___arxiv_org_abs_2501_06265
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AgoraSpeech: A multi-annotated comprehensive dataset of political discourse through the lens of humans and AI
Sermpezis, Pavlos
Karamanidis, Stelios
Paraschou, Eva
Dimitriadis, Ilias
Yfantidou, Sofia
Kouskouveli, Filitsa-Ioanna
Troboukis, Thanasis
Kiki, Kelly
Galanopoulos, Antonis
Vakali, Athena
Computation and Language
Political discourse datasets are important for gaining political insights, analyzing communication strategies or social science phenomena. Although numerous political discourse corpora exist, comprehensive, high-quality, annotated datasets are scarce. This is largely due to the substantial manual effort, multidisciplinarity, and expertise required for the nuanced annotation of rhetorical strategies and ideological contexts. In this paper, we present AgoraSpeech, a meticulously curated, high-quality dataset of 171 political speeches from six parties during the Greek national elections in 2023. The dataset includes annotations (per paragraph) for six natural language processing (NLP) tasks: text classification, topic identification, sentiment analysis, named entity recognition, polarization and populism detection. A two-step annotation was employed, starting with ChatGPT-generated annotations and followed by exhaustive human-in-the-loop validation. The dataset was initially used in a case study to provide insights during the pre-election period. However, it has general applicability by serving as a rich source of information for political and social scientists, journalists, or data scientists, while it can be used for benchmarking and fine-tuning NLP and large language models (LLMs).
title AgoraSpeech: A multi-annotated comprehensive dataset of political discourse through the lens of humans and AI
topic Computation and Language
url https://arxiv.org/abs/2501.06265