From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Seokhee, Kim, Sunkyoung, Son, Guijin, Kim, Soyeon, Hong, Yeonjung, Lee, Jinsik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
KMMLU: Measuring Massive Multitask Language Understanding in Korean
von: Son, Guijin, et al.
Veröffentlicht: (2024)
von: Son, Guijin, et al.
Veröffentlicht: (2024)
Cross-lingual QA: A Key to Unlocking In-context Cross-lingual Performance
von: Kim, Sunkyoung, et al.
Veröffentlicht: (2023)
von: Kim, Sunkyoung, et al.
Veröffentlicht: (2023)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
von: Lee, Hanwool, et al.
Veröffentlicht: (2025)
von: Lee, Hanwool, et al.
Veröffentlicht: (2025)
Revisiting the UID Hypothesis in LLM Reasoning Traces
von: Gwak, Minju, et al.
Veröffentlicht: (2025)
von: Gwak, Minju, et al.
Veröffentlicht: (2025)
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
von: Gwak, Minju, et al.
Veröffentlicht: (2025)
von: Gwak, Minju, et al.
Veröffentlicht: (2025)
KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding
von: Hwang, Bokwang, et al.
Veröffentlicht: (2025)
von: Hwang, Bokwang, et al.
Veröffentlicht: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
von: Park, Chanjun, et al.
Veröffentlicht: (2024)
BenchPreS: A Benchmark for Context-Aware Personalized Preference Selectivity of Persistent-Memory LLMs
von: Yoon, Sangyeon, et al.
Veröffentlicht: (2026)
von: Yoon, Sangyeon, et al.
Veröffentlicht: (2026)
EXAONE 3.0 7.8B Instruction Tuned Language Model
von: An, Soyoung, et al.
Veröffentlicht: (2024)
von: An, Soyoung, et al.
Veröffentlicht: (2024)
Reasoning Models Better Express Their Confidence
von: Yoon, Dongkeun, et al.
Veröffentlicht: (2025)
von: Yoon, Dongkeun, et al.
Veröffentlicht: (2025)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
von: Kim, Hyeonwoo, et al.
Veröffentlicht: (2024)
von: Kim, Hyeonwoo, et al.
Veröffentlicht: (2024)
EXAONE Deep: Reasoning Enhanced Language Models
von: Bae, Kyunghoon, et al.
Veröffentlicht: (2025)
von: Bae, Kyunghoon, et al.
Veröffentlicht: (2025)
KoBBQ: Korean Bias Benchmark for Question Answering
von: Jin, Jiho, et al.
Veröffentlicht: (2023)
von: Jin, Jiho, et al.
Veröffentlicht: (2023)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
von: Kim, Haechan, et al.
Veröffentlicht: (2026)
von: Kim, Haechan, et al.
Veröffentlicht: (2026)
KoBALT: Korean Benchmark For Advanced Linguistic Tasks
von: Shin, Hyopil, et al.
Veröffentlicht: (2025)
von: Shin, Hyopil, et al.
Veröffentlicht: (2025)
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
von: Kim, Kyuhee, et al.
Veröffentlicht: (2025)
von: Kim, Kyuhee, et al.
Veröffentlicht: (2025)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
On the Robustness of Reward Models for Language Model Alignment
von: Hong, Jiwoo, et al.
Veröffentlicht: (2025)
von: Hong, Jiwoo, et al.
Veröffentlicht: (2025)
EXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modes
von: Bae, Kyunghoon, et al.
Veröffentlicht: (2025)
von: Bae, Kyunghoon, et al.
Veröffentlicht: (2025)
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
SLM-Based Agentic AI with P-C-G: Optimized for Korean Tool Use
von: Jeon, Changhyun, et al.
Veröffentlicht: (2025)
von: Jeon, Changhyun, et al.
Veröffentlicht: (2025)
M-Prometheus: A Suite of Open Multilingual LLM Judges
von: Pombal, José, et al.
Veröffentlicht: (2025)
von: Pombal, José, et al.
Veröffentlicht: (2025)
KOMBO: Korean Character Representations Based on the Combination Rules of Subcharacters
von: Kim, SungHo, et al.
Veröffentlicht: (2026)
von: Kim, SungHo, et al.
Veröffentlicht: (2026)
KoACD: The First Korean Adolescent Dataset for Cognitive Distortion Analysis via Role-Switching Multi-LLM Negotiation
von: Kim, JunSeo, et al.
Veröffentlicht: (2025)
von: Kim, JunSeo, et al.
Veröffentlicht: (2025)
KatFishNet: Detecting LLM-Generated Korean Text through Linguistic Feature Analysis
von: Park, Shinwoo, et al.
Veröffentlicht: (2025)
von: Park, Shinwoo, et al.
Veröffentlicht: (2025)
KoCoSa: Korean Context-aware Sarcasm Detection Dataset
von: Kim, Yumin, et al.
Veröffentlicht: (2024)
von: Kim, Yumin, et al.
Veröffentlicht: (2024)
KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language
von: Kim, Yoonshik, et al.
Veröffentlicht: (2025)
von: Kim, Yoonshik, et al.
Veröffentlicht: (2025)
Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks
von: Warner, Benjamin, et al.
Veröffentlicht: (2026)
von: Warner, Benjamin, et al.
Veröffentlicht: (2026)
Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean
von: Kim, SungHo, et al.
Veröffentlicht: (2025)
von: Kim, SungHo, et al.
Veröffentlicht: (2025)
LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation
von: Zhu, Liya, et al.
Veröffentlicht: (2025)
von: Zhu, Liya, et al.
Veröffentlicht: (2025)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
von: Son, Guijin, et al.
Veröffentlicht: (2023)
von: Son, Guijin, et al.
Veröffentlicht: (2023)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
von: Son, Guijin, et al.
Veröffentlicht: (2024)
von: Son, Guijin, et al.
Veröffentlicht: (2024)
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
K-EXAONE Technical Report
von: Choi, Eunbi, et al.
Veröffentlicht: (2026)
von: Choi, Eunbi, et al.
Veröffentlicht: (2026)
Evaluating Multimodal Generative AI with Korean Educational Standards
von: Park, Sanghee, et al.
Veröffentlicht: (2025)
von: Park, Sanghee, et al.
Veröffentlicht: (2025)
DeFrame: Debiasing Large Language Models Against Framing Effects
von: Lim, Kahee, et al.
Veröffentlicht: (2026)
von: Lim, Kahee, et al.
Veröffentlicht: (2026)
Evaluating Consistencies in LLM responses through a Semantic Clustering of Question Answering
von: Lee, Yanggyu, et al.
Veröffentlicht: (2024)
von: Lee, Yanggyu, et al.
Veröffentlicht: (2024)
CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics
von: Somasekharan, Nithin, et al.
Veröffentlicht: (2025)
von: Somasekharan, Nithin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
KMMLU: Measuring Massive Multitask Language Understanding in Korean
von: Son, Guijin, et al.
Veröffentlicht: (2024) -
Cross-lingual QA: A Key to Unlocking In-context Cross-lingual Performance
von: Kim, Sunkyoung, et al.
Veröffentlicht: (2023) -
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2025) -
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
von: Lee, Hanwool, et al.
Veröffentlicht: (2025) -
Revisiting the UID Hypothesis in LLM Reasoning Traces
von: Gwak, Minju, et al.
Veröffentlicht: (2025)