From Articles to Premises: Building PrimeFacts, an Extraction Methodology and Resource for Fact-Checking Evidence

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sahitaj, Premtim, Kolanowski, Jawan, Sahitaj, Ariana, Solopova, Veronika, Upravitelev, Max, Röder, Daniel, Maab, Iffat, Yamagishi, Junichi, Möller, Sebastian, Schmitt, Vera
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917467755905024
author Sahitaj, Premtim
Kolanowski, Jawan
Sahitaj, Ariana
Solopova, Veronika
Upravitelev, Max
Röder, Daniel
Maab, Iffat
Yamagishi, Junichi
Möller, Sebastian
Schmitt, Vera
author_facet Sahitaj, Premtim
Kolanowski, Jawan
Sahitaj, Ariana
Solopova, Veronika
Upravitelev, Max
Röder, Daniel
Maab, Iffat
Yamagishi, Junichi
Möller, Sebastian
Schmitt, Vera
contents Fact-checking articles encode rich supporting evidence and reasoning, yet this evidence remains largely inaccessible to automated verification systems due to unstructured presentation. We introduce PrimeFacts, a methodology and resource for extracting fine-grained evidence from full fact-checking articles. We compile 13,106 PolitiFact articles with claims, verdicts, and all referenced sources, and we identify 49,718 in-article hyperlinks as natural anchors to pinpoint key evidence. Our framework leverages large language models (LLMs) to rewrite these anchor sentences into stand-alone, context-independent premises and investigates the extraction of additional implicit evidence. In evaluations on cross-article evidence retrieval and claim verification, the extracted premises substantially improve performance. Decontextualized evidence yields higher retrievability, achieving up to a 30 percent relative gain in Mean Reciprocal Rank over verbatim sentences, and using the evidence for verdict prediction raises Macro-F1 by 10-20 points over the baseline. These gains are consistent across different verdict granularities (2-class vs. 5-class) and model architectures. A qualitative analysis indicates that the decontextualized premises remain faithful to the original sources. Our work highlights the promise of reusing fact-checkers' evidence for automation and provides a large-scale resource of structured evidence from real-world fact-checks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_06006
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From Articles to Premises: Building PrimeFacts, an Extraction Methodology and Resource for Fact-Checking Evidence
Sahitaj, Premtim
Kolanowski, Jawan
Sahitaj, Ariana
Solopova, Veronika
Upravitelev, Max
Röder, Daniel
Maab, Iffat
Yamagishi, Junichi
Möller, Sebastian
Schmitt, Vera
Computation and Language
Fact-checking articles encode rich supporting evidence and reasoning, yet this evidence remains largely inaccessible to automated verification systems due to unstructured presentation. We introduce PrimeFacts, a methodology and resource for extracting fine-grained evidence from full fact-checking articles. We compile 13,106 PolitiFact articles with claims, verdicts, and all referenced sources, and we identify 49,718 in-article hyperlinks as natural anchors to pinpoint key evidence. Our framework leverages large language models (LLMs) to rewrite these anchor sentences into stand-alone, context-independent premises and investigates the extraction of additional implicit evidence. In evaluations on cross-article evidence retrieval and claim verification, the extracted premises substantially improve performance. Decontextualized evidence yields higher retrievability, achieving up to a 30 percent relative gain in Mean Reciprocal Rank over verbatim sentences, and using the evidence for verdict prediction raises Macro-F1 by 10-20 points over the baseline. These gains are consistent across different verdict granularities (2-class vs. 5-class) and model architectures. A qualitative analysis indicates that the decontextualized premises remain faithful to the original sources. Our work highlights the promise of reusing fact-checkers' evidence for automation and provides a large-scale resource of structured evidence from real-world fact-checks.
title From Articles to Premises: Building PrimeFacts, an Extraction Methodology and Resource for Fact-Checking Evidence
topic Computation and Language
url https://arxiv.org/abs/2605.06006