Assessing Confidence in AI-Assisted Grading of Physics Exams through Psychometrics: An Exploratory Study

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kortemeyer, Gerd, Nöhl, Julian
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908306391433216
author Kortemeyer, Gerd
Nöhl, Julian
author_facet Kortemeyer, Gerd
Nöhl, Julian
contents This study explores the use of artificial intelligence in grading high-stakes physics exams, emphasizing the application of psychometric methods, particularly Item Response Theory (IRT), to evaluate the reliability of AI-assisted grading. We examine how grading rubrics can be iteratively refined and how threshold parameters can determine when AI-generated grades are reliable versus when human intervention is necessary. By adjusting thresholds for correctness measures and uncertainty, AI can grade with high precision, significantly reducing grading workloads while maintaining accuracy. Our findings show that AI can achieve a coefficient of determination of $R^2\approx 0.91$ when handling half of the grading load, and $R^2 \approx 0.96$ for one-fifth of the load. These results demonstrate AI's potential to assist in grading large-scale assessments, reducing both human effort and associated costs. However, the study underscores the importance of human oversight in cases of uncertainty or complex problem-solving, ensuring the integrity of the grading process.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19409
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Assessing Confidence in AI-Assisted Grading of Physics Exams through Psychometrics: An Exploratory Study
Kortemeyer, Gerd
Nöhl, Julian
Physics Education
This study explores the use of artificial intelligence in grading high-stakes physics exams, emphasizing the application of psychometric methods, particularly Item Response Theory (IRT), to evaluate the reliability of AI-assisted grading. We examine how grading rubrics can be iteratively refined and how threshold parameters can determine when AI-generated grades are reliable versus when human intervention is necessary. By adjusting thresholds for correctness measures and uncertainty, AI can grade with high precision, significantly reducing grading workloads while maintaining accuracy. Our findings show that AI can achieve a coefficient of determination of $R^2\approx 0.91$ when handling half of the grading load, and $R^2 \approx 0.96$ for one-fifth of the load. These results demonstrate AI's potential to assist in grading large-scale assessments, reducing both human effort and associated costs. However, the study underscores the importance of human oversight in cases of uncertainty or complex problem-solving, ensuring the integrity of the grading process.
title Assessing Confidence in AI-Assisted Grading of Physics Exams through Psychometrics: An Exploratory Study
topic Physics Education
url https://arxiv.org/abs/2410.19409