CHiL(L)Grader: Calibrated Human-in-the-Loop Short-Answer Grading

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Raikote, Pranav, Randl, Korbinian, Miliou, Ioanna, Lakes, Athanasios, Papapetrou, Panagiotis
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908882005131264
author Raikote, Pranav
Randl, Korbinian
Miliou, Ioanna
Lakes, Athanasios
Papapetrou, Panagiotis
author_facet Raikote, Pranav
Randl, Korbinian
Miliou, Ioanna
Lakes, Athanasios
Papapetrou, Panagiotis
contents Scaling educational assessment with large language models requires not just accuracy, but the ability to recognize when predictions are trustworthy. Instruction-tuned models tend to be overconfident, and their reliability deteriorates as curricula evolve, making fully autonomous deployment unsafe in high-stakes settings. We introduce CHiL(L)Grader, the first automated grading framework that incorporates calibrated confidence estimation into a human-in-the-loop workflow. Using post-hoc temperature scaling, confidence-based selective prediction, and continual learning, CHiL(L)Grader automates only high-confidence predictions while routing uncertain cases to human graders, and adapts to evolving rubrics and unseen questions. Across three short-answer grading datasets, CHiL(L)Grader automatically scores 35-65% of responses at expert-level quality (QWK >= 0.80). A QWK gap of 0.347 between accepted and rejected predictions confirms the effectiveness of the confidence-based routing. Each correction cycle strengthens the model's grading capability as it learns from teacher feedback. These results show that uncertainty quantification is key for reliable AI-assisted grading.
format Preprint
id arxiv_https___arxiv_org_abs_2603_11957
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CHiL(L)Grader: Calibrated Human-in-the-Loop Short-Answer Grading
Raikote, Pranav
Randl, Korbinian
Miliou, Ioanna
Lakes, Athanasios
Papapetrou, Panagiotis
Computation and Language
Scaling educational assessment with large language models requires not just accuracy, but the ability to recognize when predictions are trustworthy. Instruction-tuned models tend to be overconfident, and their reliability deteriorates as curricula evolve, making fully autonomous deployment unsafe in high-stakes settings. We introduce CHiL(L)Grader, the first automated grading framework that incorporates calibrated confidence estimation into a human-in-the-loop workflow. Using post-hoc temperature scaling, confidence-based selective prediction, and continual learning, CHiL(L)Grader automates only high-confidence predictions while routing uncertain cases to human graders, and adapts to evolving rubrics and unseen questions. Across three short-answer grading datasets, CHiL(L)Grader automatically scores 35-65% of responses at expert-level quality (QWK >= 0.80). A QWK gap of 0.347 between accepted and rejected predictions confirms the effectiveness of the confidence-based routing. Each correction cycle strengthens the model's grading capability as it learns from teacher feedback. These results show that uncertainty quantification is key for reliable AI-assisted grading.
title CHiL(L)Grader: Calibrated Human-in-the-Loop Short-Answer Grading
topic Computation and Language
url https://arxiv.org/abs/2603.11957