Language Bias in Information Retrieval: The Nature of the Beast and Mitigation Methods

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Jinrui, Jiang, Fan, Baldwin, Timothy
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915484018933760
author Yang, Jinrui
Jiang, Fan
Baldwin, Timothy
author_facet Yang, Jinrui
Jiang, Fan
Baldwin, Timothy
contents Language fairness in multilingual information retrieval (MLIR) systems is crucial for ensuring equitable access to information across diverse languages. This paper sheds light on the issue, based on the assumption that queries in different languages, but with identical semantics, should yield equivalent ranking lists when retrieving on the same multilingual documents. We evaluate the degree of fairness using both traditional retrieval methods, and a DPR neural ranker based on mBERT and XLM-R. Additionally, we introduce `LaKDA', a novel loss designed to mitigate language biases in neural MLIR approaches. Our analysis exposes intrinsic language biases in current MLIR technologies, with notable disparities across the retrieval methods, and the effectiveness of LaKDA in enhancing language fairness.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06195
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language Bias in Information Retrieval: The Nature of the Beast and Mitigation Methods
Yang, Jinrui
Jiang, Fan
Baldwin, Timothy
Information Retrieval
Artificial Intelligence
Computation and Language
Language fairness in multilingual information retrieval (MLIR) systems is crucial for ensuring equitable access to information across diverse languages. This paper sheds light on the issue, based on the assumption that queries in different languages, but with identical semantics, should yield equivalent ranking lists when retrieving on the same multilingual documents. We evaluate the degree of fairness using both traditional retrieval methods, and a DPR neural ranker based on mBERT and XLM-R. Additionally, we introduce `LaKDA', a novel loss designed to mitigate language biases in neural MLIR approaches. Our analysis exposes intrinsic language biases in current MLIR technologies, with notable disparities across the retrieval methods, and the effectiveness of LaKDA in enhancing language fairness.
title Language Bias in Information Retrieval: The Nature of the Beast and Mitigation Methods
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2509.06195