Semi-automated Fact-checking in Portuguese: Corpora Enrichment using Retrieval with Claim extraction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gomes, Juliana Resplande Sant'anna, Filho, Arlindo Rodrigues Galvão
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911100213133312
author Gomes, Juliana Resplande Sant'anna
Filho, Arlindo Rodrigues Galvão
author_facet Gomes, Juliana Resplande Sant'anna
Filho, Arlindo Rodrigues Galvão
contents The accelerated dissemination of disinformation often outpaces the capacity for manual fact-checking, highlighting the urgent need for Semi-Automated Fact-Checking (SAFC) systems. Within the Portuguese language context, there is a noted scarcity of publicly available datasets that integrate external evidence, an essential component for developing robust AFC systems, as many existing resources focus solely on classification based on intrinsic text features. This dissertation addresses this gap by developing, applying, and analyzing a methodology to enrich Portuguese news corpora (Fake.Br, COVID19.BR, MuMiN-PT) with external evidence. The approach simulates a user's verification process, employing Large Language Models (LLMs, specifically Gemini 1.5 Flash) to extract the main claim from texts and search engine APIs (Google Search API, Google FactCheck Claims Search API) to retrieve relevant external documents (evidence). Additionally, a data validation and preprocessing framework, including near-duplicate detection, is introduced to enhance the quality of the base corpora.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06495
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semi-automated Fact-checking in Portuguese: Corpora Enrichment using Retrieval with Claim extraction
Gomes, Juliana Resplande Sant'anna
Filho, Arlindo Rodrigues Galvão
Computation and Language
Artificial Intelligence
Information Retrieval
The accelerated dissemination of disinformation often outpaces the capacity for manual fact-checking, highlighting the urgent need for Semi-Automated Fact-Checking (SAFC) systems. Within the Portuguese language context, there is a noted scarcity of publicly available datasets that integrate external evidence, an essential component for developing robust AFC systems, as many existing resources focus solely on classification based on intrinsic text features. This dissertation addresses this gap by developing, applying, and analyzing a methodology to enrich Portuguese news corpora (Fake.Br, COVID19.BR, MuMiN-PT) with external evidence. The approach simulates a user's verification process, employing Large Language Models (LLMs, specifically Gemini 1.5 Flash) to extract the main claim from texts and search engine APIs (Google Search API, Google FactCheck Claims Search API) to retrieve relevant external documents (evidence). Additionally, a data validation and preprocessing framework, including near-duplicate detection, is introduced to enhance the quality of the base corpora.
title Semi-automated Fact-checking in Portuguese: Corpora Enrichment using Retrieval with Claim extraction
topic Computation and Language
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2508.06495