Aligning Large Language Models to Low-Resource Languages through LLM-Based Selective Translation: A Systematic Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paul, Rakesh, Kamath, Anusha, Singla, Kanishk, Joshi, Raviraj, Vaidya, Utkarsh, Chauhan, Sanjay Singh, Wartikar, Niranjan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908595083280384
author Paul, Rakesh
Kamath, Anusha
Singla, Kanishk
Joshi, Raviraj
Vaidya, Utkarsh
Chauhan, Sanjay Singh
Wartikar, Niranjan
author_facet Paul, Rakesh
Kamath, Anusha
Singla, Kanishk
Joshi, Raviraj
Vaidya, Utkarsh
Chauhan, Sanjay Singh
Wartikar, Niranjan
contents Multilingual large language models (LLMs) often demonstrate a performance gap between English and non-English languages, particularly in low-resource settings. Aligning these models to low-resource languages is essential yet challenging due to limited high-quality data. While English alignment datasets are readily available, curating equivalent data in other languages is expensive and time-consuming. A common workaround is to translate existing English alignment data; however, standard translation techniques often fail to preserve critical elements such as code, mathematical expressions, and structured formats like JSON. In this work, we investigate LLM-based selective translation, a technique that selectively translates only the translatable parts of a text while preserving non-translatable content and sentence structure. We conduct a systematic study to explore key questions around this approach, including its effectiveness compared to vanilla translation, the importance of filtering noisy outputs, and the benefits of mixing translated samples with original English data during alignment. Our experiments focus on the low-resource Indic language Hindi and compare translations generated by Google Cloud Translation (GCP) and Llama-3.1-405B. The results highlight the promise of selective translation as a practical and effective method for improving multilingual alignment in LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2507_14304
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aligning Large Language Models to Low-Resource Languages through LLM-Based Selective Translation: A Systematic Study
Paul, Rakesh
Kamath, Anusha
Singla, Kanishk
Joshi, Raviraj
Vaidya, Utkarsh
Chauhan, Sanjay Singh
Wartikar, Niranjan
Computation and Language
Machine Learning
Multilingual large language models (LLMs) often demonstrate a performance gap between English and non-English languages, particularly in low-resource settings. Aligning these models to low-resource languages is essential yet challenging due to limited high-quality data. While English alignment datasets are readily available, curating equivalent data in other languages is expensive and time-consuming. A common workaround is to translate existing English alignment data; however, standard translation techniques often fail to preserve critical elements such as code, mathematical expressions, and structured formats like JSON. In this work, we investigate LLM-based selective translation, a technique that selectively translates only the translatable parts of a text while preserving non-translatable content and sentence structure. We conduct a systematic study to explore key questions around this approach, including its effectiveness compared to vanilla translation, the importance of filtering noisy outputs, and the benefits of mixing translated samples with original English data during alignment. Our experiments focus on the low-resource Indic language Hindi and compare translations generated by Google Cloud Translation (GCP) and Llama-3.1-405B. The results highlight the promise of selective translation as a practical and effective method for improving multilingual alignment in LLMs.
title Aligning Large Language Models to Low-Resource Languages through LLM-Based Selective Translation: A Systematic Study
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2507.14304