Aligning LLMs for Multilingual Consistency in Enterprise Applications

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Agarwal, Amit, Meghwani, Hansa, Patel, Hitesh Laxmichand, Sheng, Tao, Ravi, Sujith, Roth, Dan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912669604249600
author Agarwal, Amit
Meghwani, Hansa
Patel, Hitesh Laxmichand
Sheng, Tao
Ravi, Sujith
Roth, Dan
author_facet Agarwal, Amit
Meghwani, Hansa
Patel, Hitesh Laxmichand
Sheng, Tao
Ravi, Sujith
Roth, Dan
contents Large language models (LLMs) remain unreliable for global enterprise applications due to substantial performance gaps between high-resource and mid/low-resource languages, driven by English-centric pretraining and internal reasoning biases. This inconsistency undermines customer experience and operational reliability in multilingual settings such as customer support, content moderation, and information retrieval. Even with advanced Retrieval-Augmented Generation (RAG) systems, we observe up to an 29% accuracy drop in non-English languages compared to English. We propose a practical, batch-wise alignment strategy for fine-tuning LLMs, leveraging semantically equivalent multilingual data in each training batch to directly align model outputs across languages. This approach improves non-English accuracy by up to 23.9% without compromising English performance, model reasoning, or retrieval quality. Our method is simple to implement, scalable, and integrates seamlessly with existing LLM training & deployment pipelines, enabling more robust and equitable multilingual AI solutions in industry.
format Preprint
id arxiv_https___arxiv_org_abs_2509_23659
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aligning LLMs for Multilingual Consistency in Enterprise Applications
Agarwal, Amit
Meghwani, Hansa
Patel, Hitesh Laxmichand
Sheng, Tao
Ravi, Sujith
Roth, Dan
Computation and Language
Artificial Intelligence
68T05, 68T50, 68Q25
I.2.7; I.5.1; I.2.8
Large language models (LLMs) remain unreliable for global enterprise applications due to substantial performance gaps between high-resource and mid/low-resource languages, driven by English-centric pretraining and internal reasoning biases. This inconsistency undermines customer experience and operational reliability in multilingual settings such as customer support, content moderation, and information retrieval. Even with advanced Retrieval-Augmented Generation (RAG) systems, we observe up to an 29% accuracy drop in non-English languages compared to English. We propose a practical, batch-wise alignment strategy for fine-tuning LLMs, leveraging semantically equivalent multilingual data in each training batch to directly align model outputs across languages. This approach improves non-English accuracy by up to 23.9% without compromising English performance, model reasoning, or retrieval quality. Our method is simple to implement, scalable, and integrates seamlessly with existing LLM training & deployment pipelines, enabling more robust and equitable multilingual AI solutions in industry.
title Aligning LLMs for Multilingual Consistency in Enterprise Applications
topic Computation and Language
Artificial Intelligence
68T05, 68T50, 68Q25
I.2.7; I.5.1; I.2.8
url https://arxiv.org/abs/2509.23659