Multi-EuP: The Multilingual European Parliament Dataset for Analysis of Bias in Information Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Jinrui, Baldwin, Timothy, Cohn, Trevor
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908522987388928
author Yang, Jinrui
Baldwin, Timothy
Cohn, Trevor
author_facet Yang, Jinrui
Baldwin, Timothy
Cohn, Trevor
contents We present Multi-EuP, a new multilingual benchmark dataset, comprising 22K multi-lingual documents collected from the European Parliament, spanning 24 languages. This dataset is designed to investigate fairness in a multilingual information retrieval (IR) context to analyze both language and demographic bias in a ranking context. It boasts an authentic multilingual corpus, featuring topics translated into all 24 languages, as well as cross-lingual relevance judgments. Furthermore, it offers rich demographic information associated with its documents, facilitating the study of demographic bias. We report the effectiveness of Multi-EuP for benchmarking both monolingual and multilingual IR. We also conduct a preliminary experiment on language bias caused by the choice of tokenization strategy.
format Preprint
id arxiv_https___arxiv_org_abs_2311_01870
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Multi-EuP: The Multilingual European Parliament Dataset for Analysis of Bias in Information Retrieval
Yang, Jinrui
Baldwin, Timothy
Cohn, Trevor
Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
We present Multi-EuP, a new multilingual benchmark dataset, comprising 22K multi-lingual documents collected from the European Parliament, spanning 24 languages. This dataset is designed to investigate fairness in a multilingual information retrieval (IR) context to analyze both language and demographic bias in a ranking context. It boasts an authentic multilingual corpus, featuring topics translated into all 24 languages, as well as cross-lingual relevance judgments. Furthermore, it offers rich demographic information associated with its documents, facilitating the study of demographic bias. We report the effectiveness of Multi-EuP for benchmarking both monolingual and multilingual IR. We also conduct a preliminary experiment on language bias caused by the choice of tokenization strategy.
title Multi-EuP: The Multilingual European Parliament Dataset for Analysis of Bias in Information Retrieval
topic Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2311.01870