LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Etori, Naome A., Lu, Kevin, Karisa, Randu, Kanepajs, Arturs
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916654991015936
author Etori, Naome A.
Lu, Kevin
Karisa, Randu
Kanepajs, Arturs
author_facet Etori, Naome A.
Lu, Kevin
Karisa, Randu
Kanepajs, Arturs
contents As large language models (LLMs) rapidly advance, evaluating their performance is critical. LLMs are trained on multilingual data, but their reasoning abilities are mainly evaluated using English datasets. Hence, robust evaluation frameworks are needed using high-quality non-English datasets, especially low-resource languages (LRLs). This study evaluates eight state-of-the-art (SOTA) LLMs on Latvian and Giriama using a Massive Multitask Language Understanding (MMLU) subset curated with native speakers for linguistic and cultural relevance. Giriama is benchmarked for the first time. Our evaluation shows that OpenAI's o1 model outperforms others across all languages, scoring 92.8% in English, 88.8% in Latvian, and 70.8% in Giriama on 0-shot tasks. Mistral-large (35.6%) and Llama-70B IT (41%) have weak performance, on both Latvian and Giriama. Our results underscore the need for localized benchmarks and human evaluations in advancing cultural AI contextualization.
format Preprint
id arxiv_https___arxiv_org_abs_2503_11911
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama
Etori, Naome A.
Lu, Kevin
Karisa, Randu
Kanepajs, Arturs
Computation and Language
As large language models (LLMs) rapidly advance, evaluating their performance is critical. LLMs are trained on multilingual data, but their reasoning abilities are mainly evaluated using English datasets. Hence, robust evaluation frameworks are needed using high-quality non-English datasets, especially low-resource languages (LRLs). This study evaluates eight state-of-the-art (SOTA) LLMs on Latvian and Giriama using a Massive Multitask Language Understanding (MMLU) subset curated with native speakers for linguistic and cultural relevance. Giriama is benchmarked for the first time. Our evaluation shows that OpenAI's o1 model outperforms others across all languages, scoring 92.8% in English, 88.8% in Latvian, and 70.8% in Giriama on 0-shot tasks. Mistral-large (35.6%) and Llama-70B IT (41%) have weak performance, on both Latvian and Giriama. Our results underscore the need for localized benchmarks and human evaluations in advancing cultural AI contextualization.
title LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama
topic Computation and Language
url https://arxiv.org/abs/2503.11911