Saved in:
Bibliographic Details
Main Authors: Meyer, Gérôme, Breuer, Philip, Fürst, Jonathan
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2409.18596
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929517578158080
author Meyer, Gérôme
Breuer, Philip
Fürst, Jonathan
author_facet Meyer, Gérôme
Breuer, Philip
Fürst, Jonathan
contents Open-ended questions test a more thorough understanding than closed-ended questions and are often a preferred assessment method. However, open-ended questions are tedious to grade and subject to personal bias. Therefore, there have been efforts to speed up the grading process through automation. Short Answer Grading (SAG) systems aim to automatically score students' answers. Despite growth in SAG methods and capabilities, there exists no comprehensive short-answer grading benchmark across different subjects, grading scales, and distributions. Thus, it is hard to assess the capabilities of current automated grading methods in terms of their generalizability. In this preliminary work, we introduce the combined ASAG2024 benchmark to facilitate the comparison of automated grading systems. Combining seven commonly used short-answer grading datasets in a common structure and grading scale. For our benchmark, we evaluate a set of recent SAG methods, revealing that while LLM-based approaches reach new high scores, they still are far from reaching human performance. This opens up avenues for future research on human-machine SAG systems.
format Preprint
id arxiv_https___arxiv_org_abs_2409_18596
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ASAG2024: A Combined Benchmark for Short Answer Grading
Meyer, Gérôme
Breuer, Philip
Fürst, Jonathan
Artificial Intelligence
Computation and Language
Machine Learning
Open-ended questions test a more thorough understanding than closed-ended questions and are often a preferred assessment method. However, open-ended questions are tedious to grade and subject to personal bias. Therefore, there have been efforts to speed up the grading process through automation. Short Answer Grading (SAG) systems aim to automatically score students' answers. Despite growth in SAG methods and capabilities, there exists no comprehensive short-answer grading benchmark across different subjects, grading scales, and distributions. Thus, it is hard to assess the capabilities of current automated grading methods in terms of their generalizability. In this preliminary work, we introduce the combined ASAG2024 benchmark to facilitate the comparison of automated grading systems. Combining seven commonly used short-answer grading datasets in a common structure and grading scale. For our benchmark, we evaluate a set of recent SAG methods, revealing that while LLM-based approaches reach new high scores, they still are far from reaching human performance. This opens up avenues for future research on human-machine SAG systems.
title ASAG2024: A Combined Benchmark for Short Answer Grading
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2409.18596