LLM-as-a-Grader: Practical Insights from Large Language Model for Short-Answer and Report Evaluation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Byun, Grace, Rajwal, Swati, Choi, Jinho D.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911271165624320
author Byun, Grace
Rajwal, Swati
Choi, Jinho D.
author_facet Byun, Grace
Rajwal, Swati
Choi, Jinho D.
contents Large Language Models (LLMs) are increasingly explored for educational tasks such as grading, yet their alignment with human evaluation in real classrooms remains underexamined. In this study, we investigate the feasibility of using an LLM (GPT-4o) to evaluate short-answer quizzes and project reports in an undergraduate Computational Linguistics course. We collect responses from approximately 50 students across five quizzes and receive project reports from 14 teams. LLM-generated scores are compared against human evaluations conducted independently by the course teaching assistants (TAs). Our results show that GPT-4o achieves strong correlation with human graders (up to 0.98) and exact score agreement in 55\% of quiz cases. For project reports, it also shows strong overall alignment with human grading, while exhibiting some variability in scoring technical, open-ended responses. We release all code and sample data to support further research on LLMs in educational assessment. This work highlights both the potential and limitations of LLM-based grading systems and contributes to advancing automated grading in real-world academic settings.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10819
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLM-as-a-Grader: Practical Insights from Large Language Model for Short-Answer and Report Evaluation
Byun, Grace
Rajwal, Swati
Choi, Jinho D.
Computation and Language
Large Language Models (LLMs) are increasingly explored for educational tasks such as grading, yet their alignment with human evaluation in real classrooms remains underexamined. In this study, we investigate the feasibility of using an LLM (GPT-4o) to evaluate short-answer quizzes and project reports in an undergraduate Computational Linguistics course. We collect responses from approximately 50 students across five quizzes and receive project reports from 14 teams. LLM-generated scores are compared against human evaluations conducted independently by the course teaching assistants (TAs). Our results show that GPT-4o achieves strong correlation with human graders (up to 0.98) and exact score agreement in 55\% of quiz cases. For project reports, it also shows strong overall alignment with human grading, while exhibiting some variability in scoring technical, open-ended responses. We release all code and sample data to support further research on LLMs in educational assessment. This work highlights both the potential and limitations of LLM-based grading systems and contributes to advancing automated grading in real-world academic settings.
title LLM-as-a-Grader: Practical Insights from Large Language Model for Short-Answer and Report Evaluation
topic Computation and Language
url https://arxiv.org/abs/2511.10819