Multilingual vs Crosslingual Retrieval of Fact-Checked Claims: A Tale of Two Approaches

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ramponi, Alan, Rovera, Marco, Moro, Robert, Tonelli, Sara
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916959586615296
author Ramponi, Alan
Rovera, Marco
Moro, Robert
Tonelli, Sara
author_facet Ramponi, Alan
Rovera, Marco
Moro, Robert
Tonelli, Sara
contents Retrieval of previously fact-checked claims is a well-established task, whose automation can assist professional fact-checkers in the initial steps of information verification. Previous works have mostly tackled the task monolingually, i.e., having both the input and the retrieved claims in the same language. However, especially for languages with a limited availability of fact-checks and in case of global narratives, such as pandemics, wars, or international politics, it is crucial to be able to retrieve claims across languages. In this work, we examine strategies to improve the multilingual and crosslingual performance, namely selection of negative examples (in the supervised) and re-ranking (in the unsupervised setting). We evaluate all approaches on a dataset containing posts and claims in 47 languages (283 language combinations). We observe that the best results are obtained by using LLM-based re-ranking, followed by fine-tuning with negative examples sampled using a sentence similarity-based strategy. Most importantly, we show that crosslinguality is a setup with its own unique characteristics compared to the multilingual setup.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22118
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multilingual vs Crosslingual Retrieval of Fact-Checked Claims: A Tale of Two Approaches
Ramponi, Alan
Rovera, Marco
Moro, Robert
Tonelli, Sara
Computation and Language
Retrieval of previously fact-checked claims is a well-established task, whose automation can assist professional fact-checkers in the initial steps of information verification. Previous works have mostly tackled the task monolingually, i.e., having both the input and the retrieved claims in the same language. However, especially for languages with a limited availability of fact-checks and in case of global narratives, such as pandemics, wars, or international politics, it is crucial to be able to retrieve claims across languages. In this work, we examine strategies to improve the multilingual and crosslingual performance, namely selection of negative examples (in the supervised) and re-ranking (in the unsupervised setting). We evaluate all approaches on a dataset containing posts and claims in 47 languages (283 language combinations). We observe that the best results are obtained by using LLM-based re-ranking, followed by fine-tuning with negative examples sampled using a sentence similarity-based strategy. Most importantly, we show that crosslinguality is a setup with its own unique characteristics compared to the multilingual setup.
title Multilingual vs Crosslingual Retrieval of Fact-Checked Claims: A Tale of Two Approaches
topic Computation and Language
url https://arxiv.org/abs/2505.22118