CHiL(L)Grader: Calibrated Human-in-the-Loop Short-Answer Grading
Fuente:
arXiv
Saved in:
| Main Authors: | Raikote, Pranav, Randl, Korbinian, Miliou, Ioanna, Lakes, Athanasios, Papapetrou, Panagiotis |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Machine Learning for Early Prediction of Metastasis in a Swedish Multi-Cancer Cohort
by: Rugolon, Franco, et al.
Published: (2026)
by: Rugolon, Franco, et al.
Published: (2026)
Early prediction of the risk of ICU mortality with Deep Federated Learning
by: Randl, Korbinian, et al.
Published: (2022)
by: Randl, Korbinian, et al.
Published: (2022)
Efficient Text Classification with Conformal In-Context Learning
by: Pantelidis, Ippokratis, et al.
Published: (2025)
by: Pantelidis, Ippokratis, et al.
Published: (2025)
Evaluating the Reliability of Self-Explanations in Large Language Models
by: Randl, Korbinian, et al.
Published: (2024)
by: Randl, Korbinian, et al.
Published: (2024)
CICLe: Conformal In-Context Learning for Largescale Multi-Class Food Risk Classification
by: Randl, Korbinian, et al.
Published: (2024)
by: Randl, Korbinian, et al.
Published: (2024)
SemEval-2025 Task 9: The Food Hazard Detection Challenge
by: Randl, Korbinian, et al.
Published: (2025)
by: Randl, Korbinian, et al.
Published: (2025)
LLM-as-a-Grader: Practical Insights from Large Language Model for Short-Answer and Report Evaluation
by: Byun, Grace, et al.
Published: (2025)
by: Byun, Grace, et al.
Published: (2025)
Confidence Estimation in Automatic Short Answer Grading with LLMs
by: Cong, Longwei, et al.
Published: (2026)
by: Cong, Longwei, et al.
Published: (2026)
Grade Guard: A Smart System for Short Answer Automated Grading
by: Dadu, Niharika, et al.
Published: (2025)
by: Dadu, Niharika, et al.
Published: (2025)
RAG-E: Quantifying Retriever-Generator Alignment and Failure Modes
by: Randl, Korbinian, et al.
Published: (2026)
by: Randl, Korbinian, et al.
Published: (2026)
Enhancing LLM-Based Short Answer Grading with Retrieval-Augmented Generation
by: Chu, Yucheng, et al.
Published: (2025)
by: Chu, Yucheng, et al.
Published: (2025)
Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory
by: Cong, Longwei, et al.
Published: (2026)
by: Cong, Longwei, et al.
Published: (2026)
Enhancing Security and Strengthening Defenses in Automated Short-Answer Grading Systems
by: Yarmohammadtoosky, Sahar, et al.
Published: (2025)
by: Yarmohammadtoosky, Sahar, et al.
Published: (2025)
ASAG2024: A Combined Benchmark for Short Answer Grading
by: Meyer, Gérôme, et al.
Published: (2024)
by: Meyer, Gérôme, et al.
Published: (2024)
LLM-based Automated Grading with Human-in-the-Loop
by: Chu, Yucheng, et al.
Published: (2025)
by: Chu, Yucheng, et al.
Published: (2025)
From Flat to Structural: Enhancing Automated Short Answer Grading with GraphRAG
by: Chu, Yucheng, et al.
Published: (2026)
by: Chu, Yucheng, et al.
Published: (2026)
Enhancing Multi-Domain Automatic Short Answer Grading through an Explainable Neuro-Symbolic Pipeline
by: Künnecke, Felix, et al.
Published: (2024)
by: Künnecke, Felix, et al.
Published: (2024)
"I understand why I got this grade": Automatic Short Answer Grading with Feedback
by: Aggarwal, Dishank, et al.
Published: (2024)
by: Aggarwal, Dishank, et al.
Published: (2024)
When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
by: Ferrer, Robinson, et al.
Published: (2026)
by: Ferrer, Robinson, et al.
Published: (2026)
Can LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel Dataset
by: Henkel, Owen, et al.
Published: (2023)
by: Henkel, Owen, et al.
Published: (2023)
Automated Long Answer Grading with RiceChem Dataset
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
Statistical Comparative Analysis of Semantic Similarities and Model Transferability Across Datasets for Short Answer Grading
by: Bonthu, Sridevi, et al.
Published: (2025)
by: Bonthu, Sridevi, et al.
Published: (2025)
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs
by: Furuhashi, Momoka, et al.
Published: (2025)
by: Furuhashi, Momoka, et al.
Published: (2025)
Pensieve Grader: An AI-Powered, Ready-to-Use Platform for Effortless Handwritten STEM Grading
by: Yang, Yoonseok, et al.
Published: (2025)
by: Yang, Yoonseok, et al.
Published: (2025)
Focusing on Students, not Machines: Grounded Question Generation and Automated Answer Grading
by: Meyer, Gérôme, et al.
Published: (2025)
by: Meyer, Gérôme, et al.
Published: (2025)
Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring
by: Wu, Xuansheng, et al.
Published: (2024)
by: Wu, Xuansheng, et al.
Published: (2024)
Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
by: Wang, Yuhan, et al.
Published: (2026)
by: Wang, Yuhan, et al.
Published: (2026)
Context Over Compute Human-in-the-Loop Outperforms Iterative Chain-of-Thought Prompting in Interview Answer Quality
by: Zhu, Kewen, et al.
Published: (2026)
by: Zhu, Kewen, et al.
Published: (2026)
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
by: Fu, Jinlan, et al.
Published: (2025)
by: Fu, Jinlan, et al.
Published: (2025)
Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education
by: Henkel, Owen, et al.
Published: (2024)
by: Henkel, Owen, et al.
Published: (2024)
Reducing the Cost: Cross-Prompt Pre-Finetuning for Short Answer Scoring
by: Funayama, Hiroaki, et al.
Published: (2024)
by: Funayama, Hiroaki, et al.
Published: (2024)
Towards LLM-based Autograding for Short Textual Answers
by: Schneider, Johannes, et al.
Published: (2023)
by: Schneider, Johannes, et al.
Published: (2023)
CHiRPE: A Step Towards Real-World Clinical NLP with Clinician-Oriented Model Explanations
by: Fong, Stephanie, et al.
Published: (2026)
by: Fong, Stephanie, et al.
Published: (2026)
CHiSafetyBench: A Chinese Hierarchical Safety Benchmark for Large Language Models
by: Zhang, Wenjing, et al.
Published: (2024)
by: Zhang, Wenjing, et al.
Published: (2024)
Multi-Query Focused Disaster Summarization via Instruction-Based Prompting
by: Seeberger, Philipp, et al.
Published: (2024)
by: Seeberger, Philipp, et al.
Published: (2024)
Language Models are Few-Shot Graders
by: Zhao, Chenyan, et al.
Published: (2025)
by: Zhao, Chenyan, et al.
Published: (2025)
Large Language Models As MOOCs Graders
by: Golchin, Shahriar, et al.
Published: (2024)
by: Golchin, Shahriar, et al.
Published: (2024)
Directed Graph-alignment Approach for Identification of Gaps in Short Answers
by: Sahu, Archana, et al.
Published: (2025)
by: Sahu, Archana, et al.
Published: (2025)
Tools in the Loop: Quantifying Uncertainty of LLM Question Answering Systems That Use Tools
by: Lymperopoulos, Panagiotis, et al.
Published: (2025)
by: Lymperopoulos, Panagiotis, et al.
Published: (2025)
Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systems
by: Huang, Tianyi, et al.
Published: (2026)
by: Huang, Tianyi, et al.
Published: (2026)
Similar Items
-
Multimodal Machine Learning for Early Prediction of Metastasis in a Swedish Multi-Cancer Cohort
by: Rugolon, Franco, et al.
Published: (2026) -
Early prediction of the risk of ICU mortality with Deep Federated Learning
by: Randl, Korbinian, et al.
Published: (2022) -
Efficient Text Classification with Conformal In-Context Learning
by: Pantelidis, Ippokratis, et al.
Published: (2025) -
Evaluating the Reliability of Self-Explanations in Large Language Models
by: Randl, Korbinian, et al.
Published: (2024) -
CICLe: Conformal In-Context Learning for Largescale Multi-Class Food Risk Classification
by: Randl, Korbinian, et al.
Published: (2024)