EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908749987315712 |
|---|---|
| author | Xu, Bin Bai, Yu Sun, Huashan Lin, Yiguan Liu, Siming Liang, Xinyue Li, Yaolin Dong, Zhuangzhi Zhang, Jingren Deng, Yufan Zou, Xinyu Gao, Yang Huang, Heyan |
| author_facet | Xu, Bin Bai, Yu Sun, Huashan Lin, Yiguan Liu, Siming Liang, Xinyue Li, Yaolin Dong, Zhuangzhi Zhang, Jingren Deng, Yufan Zou, Xinyu Gao, Yang Huang, Heyan |
| contents | As large language models continue to advance, their application in educational contexts remains underexplored and under-optimized. In this paper, we address this gap by introducing the first diverse benchmark tailored for educational scenarios, incorporating synthetic data containing 9 major scenarios and over 4,000 distinct educational contexts. To enable comprehensive assessment, we propose a set of multi-dimensional evaluation metrics that cover 12 critical aspects relevant to both teachers and students. We further apply human annotation to ensure the effectiveness of the model-generated evaluation responses. Additionally, we succeed to train a relatively small-scale model on our constructed dataset and demonstrate that it can achieve performance comparable to state-of-the-art large models (e.g., Deepseek V3, Qwen Max) on the test set. Overall, this work provides a practical foundation for the development and evaluation of education-oriented language models. Code and data are released at https://github.com/ybai-nlp/EduBench. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_16160 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios Xu, Bin Bai, Yu Sun, Huashan Lin, Yiguan Liu, Siming Liang, Xinyue Li, Yaolin Dong, Zhuangzhi Zhang, Jingren Deng, Yufan Zou, Xinyu Gao, Yang Huang, Heyan Computation and Language As large language models continue to advance, their application in educational contexts remains underexplored and under-optimized. In this paper, we address this gap by introducing the first diverse benchmark tailored for educational scenarios, incorporating synthetic data containing 9 major scenarios and over 4,000 distinct educational contexts. To enable comprehensive assessment, we propose a set of multi-dimensional evaluation metrics that cover 12 critical aspects relevant to both teachers and students. We further apply human annotation to ensure the effectiveness of the model-generated evaluation responses. Additionally, we succeed to train a relatively small-scale model on our constructed dataset and demonstrate that it can achieve performance comparable to state-of-the-art large models (e.g., Deepseek V3, Qwen Max) on the test set. Overall, this work provides a practical foundation for the development and evaluation of education-oriented language models. Code and data are released at https://github.com/ybai-nlp/EduBench. |
| title | EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2505.16160 |