Language Specific Knowledge: Do Models Know Better in X than in English?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Agarwal, Ishika, Bozdag, Nimet Beyza, Patel, Nisval, Hakkani-Tür, Dilek
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918464701071360
author Agarwal, Ishika
Bozdag, Nimet Beyza
Patel, Nisval
Hakkani-Tür, Dilek
author_facet Agarwal, Ishika
Bozdag, Nimet Beyza
Patel, Nisval
Hakkani-Tür, Dilek
contents Often, multilingual language models are trained with the objective to map semantically similar content (in different languages) in the same latent space. In this paper, we show a nuance in this training objective, and find that by changing the language of the input query, we can improve the question answering ability of language models. We make two main contributions. First, we introduce the term Language Specific Knowledge (LSK) to denote queries that are best answered in an ``expert language'' for a given LLM, thereby enhancing its question-answering ability. We introduce the problem of language selection -- for some queries, language models can perform better when queried in languages other than English, sometimes even better in low-resource languages -- and the goal is to select the optimal language for the query. Second, we introduce a variety of simple to strong baselines to empirically motivate the language selection problem (including one of our own methods called LSKExtractor). During our evaluation, we employ three datasets that contain knowledge about both cultural and social behavioral norms. Overall, the results show that principled language selection can improve the performance of a language model, and that the expected question-to-language map is not always intuitive: Gemma models know most about China and Middle East in Spanish; Qwen models know most about authority and responsibility in Arabic and Chinese. Broadly, our research contributes to the open-source development of language models that are inclusive and more aligned with the cultural and linguistic contexts in which they are deployed.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14990
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Language Specific Knowledge: Do Models Know Better in X than in English?
Agarwal, Ishika
Bozdag, Nimet Beyza
Patel, Nisval
Hakkani-Tür, Dilek
Computation and Language
Often, multilingual language models are trained with the objective to map semantically similar content (in different languages) in the same latent space. In this paper, we show a nuance in this training objective, and find that by changing the language of the input query, we can improve the question answering ability of language models. We make two main contributions. First, we introduce the term Language Specific Knowledge (LSK) to denote queries that are best answered in an ``expert language'' for a given LLM, thereby enhancing its question-answering ability. We introduce the problem of language selection -- for some queries, language models can perform better when queried in languages other than English, sometimes even better in low-resource languages -- and the goal is to select the optimal language for the query. Second, we introduce a variety of simple to strong baselines to empirically motivate the language selection problem (including one of our own methods called LSKExtractor). During our evaluation, we employ three datasets that contain knowledge about both cultural and social behavioral norms. Overall, the results show that principled language selection can improve the performance of a language model, and that the expected question-to-language map is not always intuitive: Gemma models know most about China and Middle East in Spanish; Qwen models know most about authority and responsibility in Arabic and Chinese. Broadly, our research contributes to the open-source development of language models that are inclusive and more aligned with the cultural and linguistic contexts in which they are deployed.
title Language Specific Knowledge: Do Models Know Better in X than in English?
topic Computation and Language
url https://arxiv.org/abs/2505.14990