Examining the Metrics for Document-Level Claim Extraction in Czech and Slovak

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Makaiova, Lucia, Fajcik, Martin, Jarolim, Antonin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908704730775552
author Makaiova, Lucia
Fajcik, Martin
Jarolim, Antonin
author_facet Makaiova, Lucia
Fajcik, Martin
Jarolim, Antonin
contents Document-level claim extraction remains an open challenge in the field of fact-checking, and subsequently, methods for evaluating extracted claims have received limited attention. In this work, we explore approaches to aligning two sets of claims pertaining to the same source document and computing their similarity through an alignment score. We investigate techniques to identify the best possible alignment and evaluation method between claim sets, with the aim of providing a reliable evaluation framework. Our approach enables comparison between model-extracted and human-annotated claim sets, serving as a metric for assessing the extraction performance of models and also as a possible measure of inter-annotator agreement. We conduct experiments on newly collected dataset-claims extracted from comments under Czech and Slovak news articles-domains that pose additional challenges due to the informal language, strong local context, and subtleties of these closely related languages. The results draw attention to the limitations of current evaluation approaches when applied to document-level claim extraction and highlight the need for more advanced methods-ones able to correctly capture semantic similarity and evaluate essential claim properties such as atomicity, checkworthiness, and decontextualization.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14566
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Examining the Metrics for Document-Level Claim Extraction in Czech and Slovak
Makaiova, Lucia
Fajcik, Martin
Jarolim, Antonin
Computation and Language
Artificial Intelligence
Document-level claim extraction remains an open challenge in the field of fact-checking, and subsequently, methods for evaluating extracted claims have received limited attention. In this work, we explore approaches to aligning two sets of claims pertaining to the same source document and computing their similarity through an alignment score. We investigate techniques to identify the best possible alignment and evaluation method between claim sets, with the aim of providing a reliable evaluation framework. Our approach enables comparison between model-extracted and human-annotated claim sets, serving as a metric for assessing the extraction performance of models and also as a possible measure of inter-annotator agreement. We conduct experiments on newly collected dataset-claims extracted from comments under Czech and Slovak news articles-domains that pose additional challenges due to the informal language, strong local context, and subtleties of these closely related languages. The results draw attention to the limitations of current evaluation approaches when applied to document-level claim extraction and highlight the need for more advanced methods-ones able to correctly capture semantic similarity and evaluate essential claim properties such as atomicity, checkworthiness, and decontextualization.
title Examining the Metrics for Document-Level Claim Extraction in Czech and Slovak
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2511.14566