How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gupta, Aman, Zhuang, Yingying, Yu, Zhou, Zhang, Ziji, Beniwal, Anurag
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912511401394176
author Gupta, Aman
Zhuang, Yingying
Yu, Zhou
Zhang, Ziji
Beniwal, Anurag
author_facet Gupta, Aman
Zhuang, Yingying
Yu, Zhou
Zhang, Ziji
Beniwal, Anurag
contents Despite advances in the multilingual capabilities of Large Language Models (LLMs), their performance varies substantially across different languages and tasks. In multilingual retrieval-augmented generation (RAG)-based systems, knowledge bases (KB) are often shared from high-resource languages (such as English) to low-resource ones, resulting in retrieved information from the KB being in a different language than the rest of the context. In such scenarios, two common practices are pre-translation to create a mono-lingual prompt and cross-lingual prompting for direct inference. However, the impact of these choices remains unclear. In this paper, we systematically evaluate the impact of different prompt translation strategies for classification tasks with RAG-enhanced LLMs in multilingual systems. Experimental results show that an optimized prompting strategy can significantly improve knowledge sharing across languages, therefore improve the performance on the downstream classification task. The findings advocate for a broader utilization of multilingual resource sharing and cross-lingual prompt optimization for non-English languages, especially the low-resource ones.
format Preprint
id arxiv_https___arxiv_org_abs_2507_22923
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting
Gupta, Aman
Zhuang, Yingying
Yu, Zhou
Zhang, Ziji
Beniwal, Anurag
Computation and Language
Artificial Intelligence
Despite advances in the multilingual capabilities of Large Language Models (LLMs), their performance varies substantially across different languages and tasks. In multilingual retrieval-augmented generation (RAG)-based systems, knowledge bases (KB) are often shared from high-resource languages (such as English) to low-resource ones, resulting in retrieved information from the KB being in a different language than the rest of the context. In such scenarios, two common practices are pre-translation to create a mono-lingual prompt and cross-lingual prompting for direct inference. However, the impact of these choices remains unclear. In this paper, we systematically evaluate the impact of different prompt translation strategies for classification tasks with RAG-enhanced LLMs in multilingual systems. Experimental results show that an optimized prompting strategy can significantly improve knowledge sharing across languages, therefore improve the performance on the downstream classification task. The findings advocate for a broader utilization of multilingual resource sharing and cross-lingual prompt optimization for non-English languages, especially the low-resource ones.
title How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.22923