LANGALIGN: Enhancing Non-English Language Models via Cross-Lingual Embedding Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Jong Myoung, Lee, Young-Jun, Choi, Ho-Jin, Jung, Sangkeun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915213622640640
author Kim, Jong Myoung
Lee, Young-Jun
Choi, Ho-Jin
Jung, Sangkeun
author_facet Kim, Jong Myoung
Lee, Young-Jun
Choi, Ho-Jin
Jung, Sangkeun
contents While Large Language Models have gained attention, many service developers still rely on embedding-based models due to practical constraints. In such cases, the quality of fine-tuning data directly impacts performance, and English datasets are often used as seed data for training non-English models. In this study, we propose LANGALIGN, which enhances target language processing by aligning English embedding vectors with those of the target language at the interface between the language model and the task header. Experiments on Korean, Japanese, and Chinese demonstrate that LANGALIGN significantly improves performance across all three languages. Additionally, we show that LANGALIGN can be applied in reverse to convert target language data into a format that an English-based model can process.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18603
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LANGALIGN: Enhancing Non-English Language Models via Cross-Lingual Embedding Alignment
Kim, Jong Myoung
Lee, Young-Jun
Choi, Ho-Jin
Jung, Sangkeun
Computation and Language
While Large Language Models have gained attention, many service developers still rely on embedding-based models due to practical constraints. In such cases, the quality of fine-tuning data directly impacts performance, and English datasets are often used as seed data for training non-English models. In this study, we propose LANGALIGN, which enhances target language processing by aligning English embedding vectors with those of the target language at the interface between the language model and the task header. Experiments on Korean, Japanese, and Chinese demonstrate that LANGALIGN significantly improves performance across all three languages. Additionally, we show that LANGALIGN can be applied in reverse to convert target language data into a format that an English-based model can process.
title LANGALIGN: Enhancing Non-English Language Models via Cross-Lingual Embedding Alignment
topic Computation and Language
url https://arxiv.org/abs/2503.18603