Staff View: :: Library Catalog

Saved in:

Bibliographic Details
Main Authors:	Sagimbayeva, Nursulu, Bahçeci, Ruveyda Betül, Weber, Ingmar
Format:	Preprint
Published:	2025
Subjects:	Computation and Language
Online Access:	https://arxiv.org/abs/2505.19191
Tags:	Add Tag No Tags, Be the first to tag this record!

_version_	1866918034338217984
author	Sagimbayeva, Nursulu Bahçeci, Ruveyda Betül Weber, Ingmar
author_facet	Sagimbayeva, Nursulu Bahçeci, Ruveyda Betül Weber, Ingmar
contents	Inconsistent political statements represent a form of misinformation. They erode public trust and pose challenges to accountability, when left unnoticed. Detecting inconsistencies automatically could support journalists in asking clarification questions, thereby helping to keep politicians accountable. We propose the Inconsistency detection task and develop a scale of inconsistency types to prompt NLP-research in this direction. To provide a resource for detecting inconsistencies in a political domain, we present a dataset of 698 human-annotated pairs of political statements with explanations of the annotators' reasoning for 237 samples. The statements mainly come from voting assistant platforms such as Wahl-O-Mat in Germany and Smartvote in Switzerland, reflecting real-world political issues. We benchmark Large Language Models (LLMs) on our dataset and show that in general, they are as good as humans at detecting inconsistencies, and might be even better than individual humans at predicting the crowd-annotated ground-truth. However, when it comes to identifying fine-grained inconsistency types, none of the model have reached the upper bound of performance (due to natural labeling variation), thus leaving room for improvement. We make our dataset and code publicly available.
format	Preprint
id	arxiv_https___arxiv_org_abs_2505_19191
institution	arXiv
publishDate	2025
record_format	arxiv
spellingShingle	Misleading through Inconsistency: A Benchmark for Political Inconsistencies Detection Sagimbayeva, Nursulu Bahçeci, Ruveyda Betül Weber, Ingmar Computation and Language Inconsistent political statements represent a form of misinformation. They erode public trust and pose challenges to accountability, when left unnoticed. Detecting inconsistencies automatically could support journalists in asking clarification questions, thereby helping to keep politicians accountable. We propose the Inconsistency detection task and develop a scale of inconsistency types to prompt NLP-research in this direction. To provide a resource for detecting inconsistencies in a political domain, we present a dataset of 698 human-annotated pairs of political statements with explanations of the annotators' reasoning for 237 samples. The statements mainly come from voting assistant platforms such as Wahl-O-Mat in Germany and Smartvote in Switzerland, reflecting real-world political issues. We benchmark Large Language Models (LLMs) on our dataset and show that in general, they are as good as humans at detecting inconsistencies, and might be even better than individual humans at predicting the crowd-annotated ground-truth. However, when it comes to identifying fine-grained inconsistency types, none of the model have reached the upper bound of performance (due to natural labeling variation), thus leaving room for improvement. We make our dataset and code publicly available.
title	Misleading through Inconsistency: A Benchmark for Political Inconsistencies Detection
topic	Computation and Language
url	https://arxiv.org/abs/2505.19191

Similar Items