Investigating Language and Retrieval Bias in Multilingual Previously Fact-Checked Claim Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vykopal, Ivan, Karamolegkou, Antonia, Kopčan, Jaroslav, Peng, Qiwei, Javůrek, Tomáš, Gregor, Michal, Šimko, Marián
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911183907323904
author Vykopal, Ivan
Karamolegkou, Antonia
Kopčan, Jaroslav
Peng, Qiwei
Javůrek, Tomáš
Gregor, Michal
Šimko, Marián
author_facet Vykopal, Ivan
Karamolegkou, Antonia
Kopčan, Jaroslav
Peng, Qiwei
Javůrek, Tomáš
Gregor, Michal
Šimko, Marián
contents Multilingual Large Language Models (LLMs) offer powerful capabilities for cross-lingual fact-checking. However, these models often exhibit language bias, performing disproportionately better on high-resource languages such as English than on low-resource counterparts. We also present and inspect a novel concept - retrieval bias, when information retrieval systems tend to favor certain information over others, leaving the retrieval process skewed. In this paper, we study language and retrieval bias in the context of Previously Fact-Checked Claim Detection (PFCD). We evaluate six open-source multilingual LLMs across 20 languages using a fully multilingual prompting strategy, leveraging the AMC-16K dataset. By translating task prompts into each language, we uncover disparities in monolingual and cross-lingual performance and identify key trends based on model family, size, and prompting strategy. Our findings highlight persistent bias in LLM behavior and offer recommendations for improving equity in multilingual fact-checking. To investigate retrieval bias, we employed multilingual embedding models and look into the frequency of retrieved claims. Our analysis reveals that certain claims are retrieved disproportionately across different posts, leading to inflated retrieval performance for popular claims while under-representing less common ones.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25138
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Investigating Language and Retrieval Bias in Multilingual Previously Fact-Checked Claim Detection
Vykopal, Ivan
Karamolegkou, Antonia
Kopčan, Jaroslav
Peng, Qiwei
Javůrek, Tomáš
Gregor, Michal
Šimko, Marián
Computation and Language
Multilingual Large Language Models (LLMs) offer powerful capabilities for cross-lingual fact-checking. However, these models often exhibit language bias, performing disproportionately better on high-resource languages such as English than on low-resource counterparts. We also present and inspect a novel concept - retrieval bias, when information retrieval systems tend to favor certain information over others, leaving the retrieval process skewed. In this paper, we study language and retrieval bias in the context of Previously Fact-Checked Claim Detection (PFCD). We evaluate six open-source multilingual LLMs across 20 languages using a fully multilingual prompting strategy, leveraging the AMC-16K dataset. By translating task prompts into each language, we uncover disparities in monolingual and cross-lingual performance and identify key trends based on model family, size, and prompting strategy. Our findings highlight persistent bias in LLM behavior and offer recommendations for improving equity in multilingual fact-checking. To investigate retrieval bias, we employed multilingual embedding models and look into the frequency of retrieved claims. Our analysis reveals that certain claims are retrieved disproportionately across different posts, leading to inflated retrieval performance for popular claims while under-representing less common ones.
title Investigating Language and Retrieval Bias in Multilingual Previously Fact-Checked Claim Detection
topic Computation and Language
url https://arxiv.org/abs/2509.25138