CodEv: An Automated Grading Framework Leveraging Large Language Models for Consistent and Constructive Feedback

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tseng, En-Qi, Huang, Pei-Cing, Hsu, Chan, Wu, Peng-Yi, Ku, Chan-Tung, Kang, Yihuang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910847277727744
author Tseng, En-Qi
Huang, Pei-Cing
Hsu, Chan
Wu, Peng-Yi
Ku, Chan-Tung
Kang, Yihuang
author_facet Tseng, En-Qi
Huang, Pei-Cing
Hsu, Chan
Wu, Peng-Yi
Ku, Chan-Tung
Kang, Yihuang
contents Grading programming assignments is crucial for guiding students to improve their programming skills and coding styles. This study presents an automated grading framework, CodEv, which leverages Large Language Models (LLMs) to provide consistent and constructive feedback. We incorporate Chain of Thought (CoT) prompting techniques to enhance the reasoning capabilities of LLMs and ensure that the grading is aligned with human evaluation. Our framework also integrates LLM ensembles to improve the accuracy and consistency of scores, along with agreement tests to deliver reliable feedback and code review comments. The results demonstrate that the framework can yield grading results comparable to human evaluators, by using smaller LLMs. Evaluation and consistency tests of the LLMs further validate our approach, confirming the reliability of the generated scores and feedback.
format Preprint
id arxiv_https___arxiv_org_abs_2501_10421
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CodEv: An Automated Grading Framework Leveraging Large Language Models for Consistent and Constructive Feedback
Tseng, En-Qi
Huang, Pei-Cing
Hsu, Chan
Wu, Peng-Yi
Ku, Chan-Tung
Kang, Yihuang
Computers and Society
Artificial Intelligence
Human-Computer Interaction
Grading programming assignments is crucial for guiding students to improve their programming skills and coding styles. This study presents an automated grading framework, CodEv, which leverages Large Language Models (LLMs) to provide consistent and constructive feedback. We incorporate Chain of Thought (CoT) prompting techniques to enhance the reasoning capabilities of LLMs and ensure that the grading is aligned with human evaluation. Our framework also integrates LLM ensembles to improve the accuracy and consistency of scores, along with agreement tests to deliver reliable feedback and code review comments. The results demonstrate that the framework can yield grading results comparable to human evaluators, by using smaller LLMs. Evaluation and consistency tests of the LLMs further validate our approach, confirming the reliability of the generated scores and feedback.
title CodEv: An Automated Grading Framework Leveraging Large Language Models for Consistent and Constructive Feedback
topic Computers and Society
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2501.10421