EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Xu, Bin, Bai, Yu, Sun, Huashan, Lin, Yiguan, Liu, Siming, Liang, Xinyue, Li, Yaolin, Dong, Zhuangzhi, Zhang, Jingren, Deng, Yufan, Zou, Xinyu, Gao, Yang, Huang, Heyan
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908749987315712
author Xu, Bin
Bai, Yu
Sun, Huashan
Lin, Yiguan
Liu, Siming
Liang, Xinyue
Li, Yaolin
Dong, Zhuangzhi
Zhang, Jingren
Deng, Yufan
Zou, Xinyu
Gao, Yang
Huang, Heyan
author_facet Xu, Bin
Bai, Yu
Sun, Huashan
Lin, Yiguan
Liu, Siming
Liang, Xinyue
Li, Yaolin
Dong, Zhuangzhi
Zhang, Jingren
Deng, Yufan
Zou, Xinyu
Gao, Yang
Huang, Heyan
contents As large language models continue to advance, their application in educational contexts remains underexplored and under-optimized. In this paper, we address this gap by introducing the first diverse benchmark tailored for educational scenarios, incorporating synthetic data containing 9 major scenarios and over 4,000 distinct educational contexts. To enable comprehensive assessment, we propose a set of multi-dimensional evaluation metrics that cover 12 critical aspects relevant to both teachers and students. We further apply human annotation to ensure the effectiveness of the model-generated evaluation responses. Additionally, we succeed to train a relatively small-scale model on our constructed dataset and demonstrate that it can achieve performance comparable to state-of-the-art large models (e.g., Deepseek V3, Qwen Max) on the test set. Overall, this work provides a practical foundation for the development and evaluation of education-oriented language models. Code and data are released at https://github.com/ybai-nlp/EduBench.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16160
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
Xu, Bin
Bai, Yu
Sun, Huashan
Lin, Yiguan
Liu, Siming
Liang, Xinyue
Li, Yaolin
Dong, Zhuangzhi
Zhang, Jingren
Deng, Yufan
Zou, Xinyu
Gao, Yang
Huang, Heyan
Computation and Language
As large language models continue to advance, their application in educational contexts remains underexplored and under-optimized. In this paper, we address this gap by introducing the first diverse benchmark tailored for educational scenarios, incorporating synthetic data containing 9 major scenarios and over 4,000 distinct educational contexts. To enable comprehensive assessment, we propose a set of multi-dimensional evaluation metrics that cover 12 critical aspects relevant to both teachers and students. We further apply human annotation to ensure the effectiveness of the model-generated evaluation responses. Additionally, we succeed to train a relatively small-scale model on our constructed dataset and demonstrate that it can achieve performance comparable to state-of-the-art large models (e.g., Deepseek V3, Qwen Max) on the test set. Overall, this work provides a practical foundation for the development and evaluation of education-oriented language models. Code and data are released at https://github.com/ybai-nlp/EduBench.
title EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
topic Computation and Language
url https://arxiv.org/abs/2505.16160