Multi-EuP: The Multilingual European Parliament Dataset for Analysis of Bias in Information Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908522987388928 |
|---|---|
| author | Yang, Jinrui Baldwin, Timothy Cohn, Trevor |
| author_facet | Yang, Jinrui Baldwin, Timothy Cohn, Trevor |
| contents | We present Multi-EuP, a new multilingual benchmark dataset, comprising 22K multi-lingual documents collected from the European Parliament, spanning 24 languages. This dataset is designed to investigate fairness in a multilingual information retrieval (IR) context to analyze both language and demographic bias in a ranking context. It boasts an authentic multilingual corpus, featuring topics translated into all 24 languages, as well as cross-lingual relevance judgments. Furthermore, it offers rich demographic information associated with its documents, facilitating the study of demographic bias. We report the effectiveness of Multi-EuP for benchmarking both monolingual and multilingual IR. We also conduct a preliminary experiment on language bias caused by the choice of tokenization strategy. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_01870 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Multi-EuP: The Multilingual European Parliament Dataset for Analysis of Bias in Information Retrieval Yang, Jinrui Baldwin, Timothy Cohn, Trevor Computation and Language Artificial Intelligence Information Retrieval Machine Learning We present Multi-EuP, a new multilingual benchmark dataset, comprising 22K multi-lingual documents collected from the European Parliament, spanning 24 languages. This dataset is designed to investigate fairness in a multilingual information retrieval (IR) context to analyze both language and demographic bias in a ranking context. It boasts an authentic multilingual corpus, featuring topics translated into all 24 languages, as well as cross-lingual relevance judgments. Furthermore, it offers rich demographic information associated with its documents, facilitating the study of demographic bias. We report the effectiveness of Multi-EuP for benchmarking both monolingual and multilingual IR. We also conduct a preliminary experiment on language bias caused by the choice of tokenization strategy. |
| title | Multi-EuP: The Multilingual European Parliament Dataset for Analysis of Bias in Information Retrieval |
| topic | Computation and Language Artificial Intelligence Information Retrieval Machine Learning |
| url | https://arxiv.org/abs/2311.01870 |