Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zheng, Weihua, Liu, Chang, Liu, Zhengyuan, Huang, Xin, Wu, Kui, Shahrin, Muhammad Huzaifah Md, Aw, Aiti, Lee, Roy Ka-Wei
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914466840444928
author Zheng, Weihua
Liu, Chang
Liu, Zhengyuan
Huang, Xin
Wu, Kui
Shahrin, Muhammad Huzaifah Md
Aw, Aiti
Lee, Roy Ka-Wei
author_facet Zheng, Weihua
Liu, Chang
Liu, Zhengyuan
Huang, Xin
Wu, Kui
Shahrin, Muhammad Huzaifah Md
Aw, Aiti
Lee, Roy Ka-Wei
contents Multilingual Large Language Models (LLMs) struggle with cross-lingual tasks due to data imbalances between high-resource and low-resource languages, as well as monolingual bias in pre-training. Existing methods, such as bilingual fine-tuning and contrastive alignment, can improve cross-lingual performance, but they often require extensive parallel data or suffer from instability. To address these challenges, we introduce a Cross-Lingual Mapping Task during the pre-training phase, which enhances cross-lingual alignment without compromising monolingual fluency. Our approach bi-directionally maps languages within the LLM embedding space, improving both language generation and comprehension. We further propose a Language Alignment Coefficient to robustly quantify cross-lingual consistency, even in limited-data scenarios. Experimental results on machine translation (MT), cross-lingual natural language understanding (CLNLU), and cross-lingual question answering (CLQA) show that our model achieves gains of up to 11.9 BLEU points in MT, 6.72 points in CLQA BERTScore-Precision, and more than 5% in CLNLU accuracy over strong multilingual baselines. These findings highlight the potential of incorporating cross-lingual objectives into pre-training to improve multilingual LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10590
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance
Zheng, Weihua
Liu, Chang
Liu, Zhengyuan
Huang, Xin
Wu, Kui
Shahrin, Muhammad Huzaifah Md
Aw, Aiti
Lee, Roy Ka-Wei
Computation and Language
Artificial Intelligence
Multilingual Large Language Models (LLMs) struggle with cross-lingual tasks due to data imbalances between high-resource and low-resource languages, as well as monolingual bias in pre-training. Existing methods, such as bilingual fine-tuning and contrastive alignment, can improve cross-lingual performance, but they often require extensive parallel data or suffer from instability. To address these challenges, we introduce a Cross-Lingual Mapping Task during the pre-training phase, which enhances cross-lingual alignment without compromising monolingual fluency. Our approach bi-directionally maps languages within the LLM embedding space, improving both language generation and comprehension. We further propose a Language Alignment Coefficient to robustly quantify cross-lingual consistency, even in limited-data scenarios. Experimental results on machine translation (MT), cross-lingual natural language understanding (CLNLU), and cross-lingual question answering (CLQA) show that our model achieves gains of up to 11.9 BLEU points in MT, 6.72 points in CLQA BERTScore-Precision, and more than 5% in CLNLU accuracy over strong multilingual baselines. These findings highlight the potential of incorporating cross-lingual objectives into pre-training to improve multilingual LLMs.
title Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.10590