Büyük Dil Modelleri için TR-MMLU Benchmarkı: Performans Değerlendirmesi, Zorluklar ve İyileştirme Fırsatları

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bayram, M. Ali, Fincan, Ali Arda, Gümüş, Ahmet Semih, Diri, Banu, Yıldırım, Savaş, Aytaş, Öner
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915449668632576
author Bayram, M. Ali
Fincan, Ali Arda
Gümüş, Ahmet Semih
Diri, Banu
Yıldırım, Savaş
Aytaş, Öner
author_facet Bayram, M. Ali
Fincan, Ali Arda
Gümüş, Ahmet Semih
Diri, Banu
Yıldırım, Savaş
Aytaş, Öner
contents Language models have made significant advancements in understanding and generating human language, achieving remarkable success in various applications. However, evaluating these models remains a challenge, particularly for resource-limited languages like Turkish. To address this issue, we introduce the Turkish MMLU (TR-MMLU) benchmark, a comprehensive evaluation framework designed to assess the linguistic and conceptual capabilities of large language models (LLMs) in Turkish. TR-MMLU is based on a meticulously curated dataset comprising 6,200 multiple-choice questions across 62 sections within the Turkish education system. This benchmark provides a standard framework for Turkish NLP research, enabling detailed analyses of LLMs' capabilities in processing Turkish text. In this study, we evaluated state-of-the-art LLMs on TR-MMLU, highlighting areas for improvement in model design. TR-MMLU sets a new standard for advancing Turkish NLP research and inspiring future innovations.
format Preprint
id arxiv_https___arxiv_org_abs_2508_13044
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Büyük Dil Modelleri için TR-MMLU Benchmarkı: Performans Değerlendirmesi, Zorluklar ve İyileştirme Fırsatları
Bayram, M. Ali
Fincan, Ali Arda
Gümüş, Ahmet Semih
Diri, Banu
Yıldırım, Savaş
Aytaş, Öner
Computation and Language
68T50
I.2.7; I.2.6
Language models have made significant advancements in understanding and generating human language, achieving remarkable success in various applications. However, evaluating these models remains a challenge, particularly for resource-limited languages like Turkish. To address this issue, we introduce the Turkish MMLU (TR-MMLU) benchmark, a comprehensive evaluation framework designed to assess the linguistic and conceptual capabilities of large language models (LLMs) in Turkish. TR-MMLU is based on a meticulously curated dataset comprising 6,200 multiple-choice questions across 62 sections within the Turkish education system. This benchmark provides a standard framework for Turkish NLP research, enabling detailed analyses of LLMs' capabilities in processing Turkish text. In this study, we evaluated state-of-the-art LLMs on TR-MMLU, highlighting areas for improvement in model design. TR-MMLU sets a new standard for advancing Turkish NLP research and inspiring future innovations.
title Büyük Dil Modelleri için TR-MMLU Benchmarkı: Performans Değerlendirmesi, Zorluklar ve İyileştirme Fırsatları
topic Computation and Language
68T50
I.2.7; I.2.6
url https://arxiv.org/abs/2508.13044