SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Pramodya, Ashmari, Nelki, Nirasha, Shalinda, Heshan, Liyanage, Chamila, Sakai, Yusuke, Pushpananda, Randil, Weerasinghe, Ruvan, Kamigaito, Hidetaka, Watanabe, Taro
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911136733986816
author Pramodya, Ashmari
Nelki, Nirasha
Shalinda, Heshan
Liyanage, Chamila
Sakai, Yusuke
Pushpananda, Randil
Weerasinghe, Ruvan
Kamigaito, Hidetaka
Watanabe, Taro
author_facet Pramodya, Ashmari
Nelki, Nirasha
Shalinda, Heshan
Liyanage, Chamila
Sakai, Yusuke
Pushpananda, Randil
Weerasinghe, Ruvan
Kamigaito, Hidetaka
Watanabe, Taro
contents Large Language Models (LLMs) demonstrate impressive general knowledge and reasoning abilities, yet their evaluation has predominantly focused on global or anglocentric subjects, often neglecting low-resource languages and culturally specific content. While recent multilingual benchmarks attempt to bridge this gap, many rely on automatic translation, which can introduce errors and misrepresent the original cultural context. To address this, we introduce SinhalaMMLU, the first multiple-choice question answering benchmark designed specifically for Sinhala, a low-resource language. The dataset includes over 7,000 questions spanning secondary to collegiate education levels, aligned with the Sri Lankan national curriculum, and covers six domains and 30 subjects, encompassing both general academic topics and culturally grounded knowledge. We evaluate 26 LLMs on SinhalaMMLU and observe that, while Claude 3.5 sonnet and GPT-4o achieve the highest average accuracies at 67% and 62% respectively, overall model performance remains limited. In particular, models struggle in culturally rich domains such as the Humanities, revealing substantial room for improvement in adapting LLMs to low-resource and culturally specific contexts.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03162
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala
Pramodya, Ashmari
Nelki, Nirasha
Shalinda, Heshan
Liyanage, Chamila
Sakai, Yusuke
Pushpananda, Randil
Weerasinghe, Ruvan
Kamigaito, Hidetaka
Watanabe, Taro
Computation and Language
Large Language Models (LLMs) demonstrate impressive general knowledge and reasoning abilities, yet their evaluation has predominantly focused on global or anglocentric subjects, often neglecting low-resource languages and culturally specific content. While recent multilingual benchmarks attempt to bridge this gap, many rely on automatic translation, which can introduce errors and misrepresent the original cultural context. To address this, we introduce SinhalaMMLU, the first multiple-choice question answering benchmark designed specifically for Sinhala, a low-resource language. The dataset includes over 7,000 questions spanning secondary to collegiate education levels, aligned with the Sri Lankan national curriculum, and covers six domains and 30 subjects, encompassing both general academic topics and culturally grounded knowledge. We evaluate 26 LLMs on SinhalaMMLU and observe that, while Claude 3.5 sonnet and GPT-4o achieve the highest average accuracies at 67% and 62% respectively, overall model performance remains limited. In particular, models struggle in culturally rich domains such as the Humanities, revealing substantial room for improvement in adapting LLMs to low-resource and culturally specific contexts.
title SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala
topic Computation and Language
url https://arxiv.org/abs/2509.03162