KMMLU: Measuring Massive Multitask Language Understanding in Korean
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Son, Guijin, Lee, Hanwool, Kim, Sungdong, Kim, Seungone, Muennighoff, Niklas, Choi, Taekyoon, Park, Cheonbok, Yoo, Kang Min, Biderman, Stella |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
von: Hong, Seokhee, et al.
Veröffentlicht: (2025)
von: Hong, Seokhee, et al.
Veröffentlicht: (2025)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
von: Son, Guijin, et al.
Veröffentlicht: (2023)
von: Son, Guijin, et al.
Veröffentlicht: (2023)
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
von: Lee, Hanwool, et al.
Veröffentlicht: (2025)
von: Lee, Hanwool, et al.
Veröffentlicht: (2025)
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
von: Choi, Dasol, et al.
Veröffentlicht: (2025)
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once?
von: Son, Guijin, et al.
Veröffentlicht: (2024)
von: Son, Guijin, et al.
Veröffentlicht: (2024)
Can Language Models Evaluate Human Written Text? Case Study on Korean Student Writing for Education
von: Kim, Seungyoon, et al.
Veröffentlicht: (2024)
von: Kim, Seungyoon, et al.
Veröffentlicht: (2024)
Multi-Step Reasoning in Korean and the Emergent Mirage
von: Son, Guijin, et al.
Veröffentlicht: (2025)
von: Son, Guijin, et al.
Veröffentlicht: (2025)
Cross-lingual Collapse: How Language-Centric Foundation Models Shape Reasoning in Large Language Models
von: Park, Cheonbok, et al.
Veröffentlicht: (2025)
von: Park, Cheonbok, et al.
Veröffentlicht: (2025)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
Aligning Language Models to Explicitly Handle Ambiguity
von: Kim, Hyuhng Joon, et al.
Veröffentlicht: (2024)
von: Kim, Hyuhng Joon, et al.
Veröffentlicht: (2024)
Enhancing Hallucination Detection via Future Context
von: Lee, Joosung, et al.
Veröffentlicht: (2025)
von: Lee, Joosung, et al.
Veröffentlicht: (2025)
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models
von: Choi, Dasol, et al.
Veröffentlicht: (2026)
von: Choi, Dasol, et al.
Veröffentlicht: (2026)
Aligning Large Language Models by On-Policy Self-Judgment
von: Lee, Sangkyu, et al.
Veröffentlicht: (2024)
von: Lee, Sangkyu, et al.
Veröffentlicht: (2024)
KAIO: A Collection of More Challenging Korean Questions
von: Lee, Nahyun, et al.
Veröffentlicht: (2025)
von: Lee, Nahyun, et al.
Veröffentlicht: (2025)
Measuring Massive Multitask Chinese Understanding
von: Zeng, Hui
Veröffentlicht: (2023)
von: Zeng, Hui
Veröffentlicht: (2023)
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
von: Ko, Hyunwoo, et al.
Veröffentlicht: (2025)
von: Ko, Hyunwoo, et al.
Veröffentlicht: (2025)
LangBridge: Multilingual Reasoning Without Multilingual Supervision
von: Yoon, Dongkeun, et al.
Veröffentlicht: (2024)
von: Yoon, Dongkeun, et al.
Veröffentlicht: (2024)
TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish
von: Yüksel, Arda, et al.
Veröffentlicht: (2024)
von: Yüksel, Arda, et al.
Veröffentlicht: (2024)
BnMMLU: Measuring Massive Multitask Language Understanding in Bengali
von: Joy, Saman Sarker, et al.
Veröffentlicht: (2025)
von: Joy, Saman Sarker, et al.
Veröffentlicht: (2025)
KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding
von: Hwang, Bokwang, et al.
Veröffentlicht: (2025)
von: Hwang, Bokwang, et al.
Veröffentlicht: (2025)
Adaptive Contrastive Decoding in Retrieval-Augmented Generation for Handling Noisy Contexts
von: Kim, Youna, et al.
Veröffentlicht: (2024)
von: Kim, Youna, et al.
Veröffentlicht: (2024)
Rethinking the Role of Proxy Rewards in Language Model Alignment
von: Kim, Sungdong, et al.
Veröffentlicht: (2024)
von: Kim, Sungdong, et al.
Veröffentlicht: (2024)
FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
von: Ye, Seonghyeon, et al.
Veröffentlicht: (2023)
von: Ye, Seonghyeon, et al.
Veröffentlicht: (2023)
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
von: Choi, Dasol, et al.
Veröffentlicht: (2024)
von: Choi, Dasol, et al.
Veröffentlicht: (2024)
Won: Establishing Best Practices for Korean Financial NLP
von: Son, Guijin, et al.
Veröffentlicht: (2025)
von: Son, Guijin, et al.
Veröffentlicht: (2025)
Mitigating Semantic Leakage in Cross-lingual Embeddings via Orthogonality Constraint
von: Ki, Dayeon, et al.
Veröffentlicht: (2024)
von: Ki, Dayeon, et al.
Veröffentlicht: (2024)
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
von: Kim, Jeonghoon, et al.
Veröffentlicht: (2025)
von: Kim, Jeonghoon, et al.
Veröffentlicht: (2025)
Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models
von: Kim, Yeeun, et al.
Veröffentlicht: (2024)
von: Kim, Yeeun, et al.
Veröffentlicht: (2024)
ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search
von: Lee, Hyunseok, et al.
Veröffentlicht: (2025)
von: Lee, Hyunseok, et al.
Veröffentlicht: (2025)
Measuring Sycophancy of Language Models in Multi-turn Dialogues
von: Hong, Jiseung, et al.
Veröffentlicht: (2025)
von: Hong, Jiseung, et al.
Veröffentlicht: (2025)
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
von: Lee, Nahyun, et al.
Veröffentlicht: (2026)
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
von: Gwak, Minju, et al.
Veröffentlicht: (2025)
von: Gwak, Minju, et al.
Veröffentlicht: (2025)
Revisiting the UID Hypothesis in LLM Reasoning Traces
von: Gwak, Minju, et al.
Veröffentlicht: (2025)
von: Gwak, Minju, et al.
Veröffentlicht: (2025)
Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
von: Kang, Hyolim, et al.
Veröffentlicht: (2025)
von: Kang, Hyolim, et al.
Veröffentlicht: (2025)
Learning Massive-scale Partial Correlation Networks in Clinical Multi-omics Studies with HP-ACCORD
von: Lee, Sungdong, et al.
Veröffentlicht: (2024)
von: Lee, Sungdong, et al.
Veröffentlicht: (2024)
Prometheus-Vision: Vision-Language Model as a Judge for Fine-Grained Evaluation
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
Genomic insights into inherited bone marrow failure syndromes in a Korean population
von: Jong‐Mi Lee, et al.
Veröffentlicht: (2024)
von: Jong‐Mi Lee, et al.
Veröffentlicht: (2024)
Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models
von: Jung, Kyudan, et al.
Veröffentlicht: (2026)
von: Jung, Kyudan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
von: Hong, Seokhee, et al.
Veröffentlicht: (2025) -
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
von: Son, Guijin, et al.
Veröffentlicht: (2023) -
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
von: Lee, Nahyun, et al.
Veröffentlicht: (2026) -
Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models
von: Lee, Hanwool, et al.
Veröffentlicht: (2025) -
Ko-PIQA: A Korean Physical Commonsense Reasoning Dataset with Cultural Context
von: Choi, Dasol, et al.
Veröffentlicht: (2025)