Not All Languages are Equal: Insights into Multilingual Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Suhang, Tang, Jialong, Yang, Baosong, Wang, Ante, Jia, Kaidi, Yu, Jiawei, Yao, Junfeng, Su, Jinsong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912093261791232
author Wu, Suhang
Tang, Jialong
Yang, Baosong
Wang, Ante
Jia, Kaidi
Yu, Jiawei
Yao, Junfeng
Su, Jinsong
author_facet Wu, Suhang
Tang, Jialong
Yang, Baosong
Wang, Ante
Jia, Kaidi
Yu, Jiawei
Yao, Junfeng
Su, Jinsong
contents RALMs (Retrieval-Augmented Language Models) broaden their knowledge scope by incorporating external textual resources. However, the multilingual nature of global knowledge necessitates RALMs to handle diverse languages, a topic that has received limited research focus. In this work, we propose \textit{Futurepedia}, a carefully crafted benchmark containing parallel texts across eight representative languages. We evaluate six multilingual RALMs using our benchmark to explore the challenges of multilingual RALMs. Experimental results reveal linguistic inequalities: 1) high-resource languages stand out in Monolingual Knowledge Extraction; 2) Indo-European languages lead RALMs to provide answers directly from documents, alleviating the challenge of expressing answers across languages; 3) English benefits from RALMs' selection bias and speaks louder in multilingual knowledge selection. Based on these findings, we offer advice for improving multilingual Retrieval Augmented Generation. For monolingual knowledge extraction, careful attention must be paid to cascading errors from translating low-resource languages into high-resource ones. In cross-lingual knowledge transfer, encouraging RALMs to provide answers within documents in different languages can improve transfer performance. For multilingual knowledge selection, incorporating more non-English documents and repositioning English documents can help mitigate RALMs' selection bias. Through comprehensive experiments, we underscore the complexities inherent in multilingual RALMs and offer valuable insights for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2410_21970
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Not All Languages are Equal: Insights into Multilingual Retrieval-Augmented Generation
Wu, Suhang
Tang, Jialong
Yang, Baosong
Wang, Ante
Jia, Kaidi
Yu, Jiawei
Yao, Junfeng
Su, Jinsong
Computation and Language
RALMs (Retrieval-Augmented Language Models) broaden their knowledge scope by incorporating external textual resources. However, the multilingual nature of global knowledge necessitates RALMs to handle diverse languages, a topic that has received limited research focus. In this work, we propose \textit{Futurepedia}, a carefully crafted benchmark containing parallel texts across eight representative languages. We evaluate six multilingual RALMs using our benchmark to explore the challenges of multilingual RALMs. Experimental results reveal linguistic inequalities: 1) high-resource languages stand out in Monolingual Knowledge Extraction; 2) Indo-European languages lead RALMs to provide answers directly from documents, alleviating the challenge of expressing answers across languages; 3) English benefits from RALMs' selection bias and speaks louder in multilingual knowledge selection. Based on these findings, we offer advice for improving multilingual Retrieval Augmented Generation. For monolingual knowledge extraction, careful attention must be paid to cascading errors from translating low-resource languages into high-resource ones. In cross-lingual knowledge transfer, encouraging RALMs to provide answers within documents in different languages can improve transfer performance. For multilingual knowledge selection, incorporating more non-English documents and repositioning English documents can help mitigate RALMs' selection bias. Through comprehensive experiments, we underscore the complexities inherent in multilingual RALMs and offer valuable insights for future research.
title Not All Languages are Equal: Insights into Multilingual Retrieval-Augmented Generation
topic Computation and Language
url https://arxiv.org/abs/2410.21970