KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kim, Dongjun, Park, Chanhee, Park, Chanjun, Lim, Heuiseok
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917021582622720
author Kim, Dongjun
Park, Chanhee
Park, Chanjun
Lim, Heuiseok
author_facet Kim, Dongjun
Park, Chanhee
Park, Chanjun
Lim, Heuiseok
contents The instruction-following capabilities of large language models (LLMs) are pivotal for numerous applications, from conversational agents to complex reasoning systems. However, current evaluations predominantly focus on English models, neglecting the linguistic and cultural nuances of other languages. Specifically, Korean, with its distinct syntax, rich morphological features, honorific system, and dual numbering systems, lacks a dedicated benchmark for assessing open-ended instruction-following capabilities. To address this gap, we introduce the Korean Instruction-following Task Evaluation (KITE), a comprehensive benchmark designed to evaluate both general and Korean-specific instructions. Unlike existing Korean benchmarks that focus mainly on factual knowledge or multiple-choice testing, KITE directly targets diverse, open-ended instruction-following tasks. Our evaluation pipeline combines automated metrics with human assessments, revealing performance disparities across models and providing deeper insights into their strengths and weaknesses. By publicly releasing the KITE dataset and code, we aim to foster further research on culturally and linguistically inclusive LLM development and inspire similar endeavors for other underrepresented languages.
format Preprint
id arxiv_https___arxiv_org_abs_2510_15558
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
Kim, Dongjun
Park, Chanhee
Park, Chanjun
Lim, Heuiseok
Computation and Language
Artificial Intelligence
The instruction-following capabilities of large language models (LLMs) are pivotal for numerous applications, from conversational agents to complex reasoning systems. However, current evaluations predominantly focus on English models, neglecting the linguistic and cultural nuances of other languages. Specifically, Korean, with its distinct syntax, rich morphological features, honorific system, and dual numbering systems, lacks a dedicated benchmark for assessing open-ended instruction-following capabilities. To address this gap, we introduce the Korean Instruction-following Task Evaluation (KITE), a comprehensive benchmark designed to evaluate both general and Korean-specific instructions. Unlike existing Korean benchmarks that focus mainly on factual knowledge or multiple-choice testing, KITE directly targets diverse, open-ended instruction-following tasks. Our evaluation pipeline combines automated metrics with human assessments, revealing performance disparities across models and providing deeper insights into their strengths and weaknesses. By publicly releasing the KITE dataset and code, we aim to foster further research on culturally and linguistically inclusive LLM development and inspire similar endeavors for other underrepresented languages.
title KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2510.15558