A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xu, Yuemei, Hu, Ling, Zhao, Jiayi, Qiu, Zihan, XU, Kexin, Ye, Yuqi, Gu, Hanwen
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917861985878016
author Xu, Yuemei
Hu, Ling
Zhao, Jiayi
Qiu, Zihan
XU, Kexin
Ye, Yuqi
Gu, Hanwen
author_facet Xu, Yuemei
Hu, Ling
Zhao, Jiayi
Qiu, Zihan
XU, Kexin
Ye, Yuqi
Gu, Hanwen
contents Based on the foundation of Large Language Models (LLMs), Multilingual LLMs (MLLMs) have been developed to address the challenges faced in multilingual natural language processing, hoping to achieve knowledge transfer from high-resource languages to low-resource languages. However, significant limitations and challenges still exist, such as language imbalance, multilingual alignment, and inherent bias. In this paper, we aim to provide a comprehensive analysis of MLLMs, delving deeply into discussions surrounding these critical issues. First of all, we start by presenting an overview of MLLMs, covering their evolutions, key techniques, and multilingual capacities. Secondly, we explore the multilingual training corpora of MLLMs and the multilingual datasets oriented for downstream tasks that are crucial to enhance the cross-lingual capability of MLLMs. Thirdly, we survey the state-of-the-art studies of multilingual representations and investigate whether the current MLLMs can learn a universal language representation. Fourthly, we discuss bias on MLLMs, including its categories, evaluation metrics, and debiasing techniques. Finally, we discuss existing challenges and point out promising research directions of MLLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2404_00929
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
Xu, Yuemei
Hu, Ling
Zhao, Jiayi
Qiu, Zihan
XU, Kexin
Ye, Yuqi
Gu, Hanwen
Computation and Language
Artificial Intelligence
Based on the foundation of Large Language Models (LLMs), Multilingual LLMs (MLLMs) have been developed to address the challenges faced in multilingual natural language processing, hoping to achieve knowledge transfer from high-resource languages to low-resource languages. However, significant limitations and challenges still exist, such as language imbalance, multilingual alignment, and inherent bias. In this paper, we aim to provide a comprehensive analysis of MLLMs, delving deeply into discussions surrounding these critical issues. First of all, we start by presenting an overview of MLLMs, covering their evolutions, key techniques, and multilingual capacities. Secondly, we explore the multilingual training corpora of MLLMs and the multilingual datasets oriented for downstream tasks that are crucial to enhance the cross-lingual capability of MLLMs. Thirdly, we survey the state-of-the-art studies of multilingual representations and investigate whether the current MLLMs can learn a universal language representation. Fourthly, we discuss bias on MLLMs, including its categories, evaluation metrics, and debiasing techniques. Finally, we discuss existing challenges and point out promising research directions of MLLMs.
title A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2404.00929