CLEAR: Cross-Lingual Enhancement in Alignment via Reverse-training

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lee, Seungyoon, Kim, Minhyuk, Hong, Seongtae, Jang, Youngjoon, Oh, Dongsuk, Lim, Heuiseok
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913028803395584
author Lee, Seungyoon
Kim, Minhyuk
Hong, Seongtae
Jang, Youngjoon
Oh, Dongsuk
Lim, Heuiseok
author_facet Lee, Seungyoon
Kim, Minhyuk
Hong, Seongtae
Jang, Youngjoon
Oh, Dongsuk
Lim, Heuiseok
contents Existing multilingual embedding models often encounter challenges in cross-lingual scenarios due to imbalanced linguistic resources and less consideration of cross-lingual alignment during training. Although standardized contrastive learning approaches for cross-lingual adaptation are widely adopted, they may struggle to capture fundamental alignment between languages and degrade performance in well-aligned languages such as English. To address these challenges, we propose Cross-Lingual Enhancement in Retrieval via Reverse-training (CLEAR), a novel loss function utilizing a reverse training scheme to improve retrieval performance across diverse cross-lingual retrieval scenarios. CLEAR leverages an English passage as a bridge to strengthen alignments between the target language and English, ensuring robust performance in the cross-lingual retrieval task. Our extensive experiments demonstrate that CLEAR achieves notable improvements in cross-lingual scenarios, with gains up to 15%, particularly in low-resource languages, while minimizing performance degradation in English. Furthermore, our findings highlight that CLEAR offers promising effectiveness even in multilingual training, suggesting its potential for broad application and scalability. We release the code at https://github.com/dltmddbs100/CLEAR.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05821
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CLEAR: Cross-Lingual Enhancement in Alignment via Reverse-training
Lee, Seungyoon
Kim, Minhyuk
Hong, Seongtae
Jang, Youngjoon
Oh, Dongsuk
Lim, Heuiseok
Computation and Language
Information Retrieval
Existing multilingual embedding models often encounter challenges in cross-lingual scenarios due to imbalanced linguistic resources and less consideration of cross-lingual alignment during training. Although standardized contrastive learning approaches for cross-lingual adaptation are widely adopted, they may struggle to capture fundamental alignment between languages and degrade performance in well-aligned languages such as English. To address these challenges, we propose Cross-Lingual Enhancement in Retrieval via Reverse-training (CLEAR), a novel loss function utilizing a reverse training scheme to improve retrieval performance across diverse cross-lingual retrieval scenarios. CLEAR leverages an English passage as a bridge to strengthen alignments between the target language and English, ensuring robust performance in the cross-lingual retrieval task. Our extensive experiments demonstrate that CLEAR achieves notable improvements in cross-lingual scenarios, with gains up to 15%, particularly in low-resource languages, while minimizing performance degradation in English. Furthermore, our findings highlight that CLEAR offers promising effectiveness even in multilingual training, suggesting its potential for broad application and scalability. We release the code at https://github.com/dltmddbs100/CLEAR.
title CLEAR: Cross-Lingual Enhancement in Alignment via Reverse-training
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2604.05821