EduEval: A Hierarchical Cognitive Benchmark for Evaluating Large Language Models in Chinese Education

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Guoqing, Zhu, Jia, Guo, Hanghui, Shi, Weijie, Cui, Yue, Shen, Jiawei, Li, Zilong, Liang, Yidan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917113697927168
author Ma, Guoqing
Zhu, Jia
Guo, Hanghui
Shi, Weijie
Cui, Yue
Shen, Jiawei
Li, Zilong
Liang, Yidan
author_facet Ma, Guoqing
Zhu, Jia
Guo, Hanghui
Shi, Weijie
Cui, Yue
Shen, Jiawei
Li, Zilong
Liang, Yidan
contents Large language models (LLMs) demonstrate significant potential for educational applications. However, their unscrutinized deployment poses risks to educational standards, underscoring the need for rigorous evaluation. We introduce EduEval, a comprehensive hierarchical benchmark for evaluating LLMs in Chinese K-12 education. This benchmark makes three key contributions: (1) Cognitive Framework: We propose the EduAbility Taxonomy, which unifies Bloom's Taxonomy and Webb's Depth of Knowledge to organize tasks across six cognitive dimensions including Memorization, Understanding, Application, Reasoning, Creativity, and Ethics. (2) Authenticity: Our benchmark integrates real exam questions, classroom conversation, student essays, and expert-designed prompts to reflect genuine educational challenges; (3) Scale: EduEval comprises 24 distinct task types with over 11,000 questions spanning primary to high school levels. We evaluate 14 leading LLMs under both zero-shot and few-shot settings, revealing that while models perform well on factual tasks, they struggle with classroom dialogue classification and exhibit inconsistent results in creative content generation. Interestingly, several open source models outperform proprietary systems on complex educational reasoning. Few-shot prompting shows varying effectiveness across cognitive dimensions, suggesting that different educational objectives require tailored approaches. These findings provide targeted benchmarking metrics for developing LLMs specifically optimized for diverse Chinese educational tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00290
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EduEval: A Hierarchical Cognitive Benchmark for Evaluating Large Language Models in Chinese Education
Ma, Guoqing
Zhu, Jia
Guo, Hanghui
Shi, Weijie
Cui, Yue
Shen, Jiawei
Li, Zilong
Liang, Yidan
Computation and Language
Artificial Intelligence
Large language models (LLMs) demonstrate significant potential for educational applications. However, their unscrutinized deployment poses risks to educational standards, underscoring the need for rigorous evaluation. We introduce EduEval, a comprehensive hierarchical benchmark for evaluating LLMs in Chinese K-12 education. This benchmark makes three key contributions: (1) Cognitive Framework: We propose the EduAbility Taxonomy, which unifies Bloom's Taxonomy and Webb's Depth of Knowledge to organize tasks across six cognitive dimensions including Memorization, Understanding, Application, Reasoning, Creativity, and Ethics. (2) Authenticity: Our benchmark integrates real exam questions, classroom conversation, student essays, and expert-designed prompts to reflect genuine educational challenges; (3) Scale: EduEval comprises 24 distinct task types with over 11,000 questions spanning primary to high school levels. We evaluate 14 leading LLMs under both zero-shot and few-shot settings, revealing that while models perform well on factual tasks, they struggle with classroom dialogue classification and exhibit inconsistent results in creative content generation. Interestingly, several open source models outperform proprietary systems on complex educational reasoning. Few-shot prompting shows varying effectiveness across cognitive dimensions, suggesting that different educational objectives require tailored approaches. These findings provide targeted benchmarking metrics for developing LLMs specifically optimized for diverse Chinese educational tasks.
title EduEval: A Hierarchical Cognitive Benchmark for Evaluating Large Language Models in Chinese Education
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.00290