MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fabbri, Alexander R., Mares, Diego, Flores, Jorge, Mankikar, Meher, Hernandez, Ernesto, Lee, Dean, Liu, Bing, Xing, Chen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911072927088640
author Fabbri, Alexander R.
Mares, Diego
Flores, Jorge
Mankikar, Meher
Hernandez, Ernesto
Lee, Dean
Liu, Bing
Xing, Chen
author_facet Fabbri, Alexander R.
Mares, Diego
Flores, Jorge
Mankikar, Meher
Hernandez, Ernesto
Lee, Dean
Liu, Bing
Xing, Chen
contents Although recent Large Language Models (LLMs) have shown rapid improvement on reasoning benchmarks in English, the evaluation of such LLMs' multilingual reasoning capability across diverse languages and cultural contexts remains limited. Existing multilingual reasoning benchmarks are typically constructed by translating existing English reasoning benchmarks, biasing these benchmarks towards reasoning problems with context in English language/cultures. In this work, we introduce the Multilingual Native Reasoning Challenge (MultiNRC), a benchmark designed to assess LLMs on more than 1,000 native, linguistic and culturally grounded reasoning questions written by native speakers in French, Spanish, and Chinese. MultiNRC covers four core reasoning categories: language-specific linguistic reasoning, wordplay & riddles, cultural/tradition reasoning, and math reasoning with cultural relevance. For cultural/tradition reasoning and math reasoning with cultural relevance, we also provide English equivalent translations of the multilingual questions by manual translation from native speakers fluent in English. This set of English equivalents can provide a direct comparison of LLM reasoning capacity in other languages vs. English on the same reasoning questions. We systematically evaluate current 14 leading LLMs covering most LLM families on MultiNRC and its English equivalent set. The results show that (1) current LLMs are still not good at native multilingual reasoning, with none scoring above 50% on MultiNRC; (2) LLMs exhibit distinct strengths and weaknesses in handling linguistic, cultural, and logical reasoning tasks; (3) Most models perform substantially better in math reasoning in English compared to in original languages (+10%), indicating persistent challenges with culturally grounded knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2507_17476
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
Fabbri, Alexander R.
Mares, Diego
Flores, Jorge
Mankikar, Meher
Hernandez, Ernesto
Lee, Dean
Liu, Bing
Xing, Chen
Computation and Language
Artificial Intelligence
Although recent Large Language Models (LLMs) have shown rapid improvement on reasoning benchmarks in English, the evaluation of such LLMs' multilingual reasoning capability across diverse languages and cultural contexts remains limited. Existing multilingual reasoning benchmarks are typically constructed by translating existing English reasoning benchmarks, biasing these benchmarks towards reasoning problems with context in English language/cultures. In this work, we introduce the Multilingual Native Reasoning Challenge (MultiNRC), a benchmark designed to assess LLMs on more than 1,000 native, linguistic and culturally grounded reasoning questions written by native speakers in French, Spanish, and Chinese. MultiNRC covers four core reasoning categories: language-specific linguistic reasoning, wordplay & riddles, cultural/tradition reasoning, and math reasoning with cultural relevance. For cultural/tradition reasoning and math reasoning with cultural relevance, we also provide English equivalent translations of the multilingual questions by manual translation from native speakers fluent in English. This set of English equivalents can provide a direct comparison of LLM reasoning capacity in other languages vs. English on the same reasoning questions. We systematically evaluate current 14 leading LLMs covering most LLM families on MultiNRC and its English equivalent set. The results show that (1) current LLMs are still not good at native multilingual reasoning, with none scoring above 50% on MultiNRC; (2) LLMs exhibit distinct strengths and weaknesses in handling linguistic, cultural, and logical reasoning tasks; (3) Most models perform substantially better in math reasoning in English compared to in original languages (+10%), indicating persistent challenges with culturally grounded knowledge.
title MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.17476