Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory
Fuente:
arXiv
Saved in:
| Main Authors: | Cong, Longwei, Hahn, Sonja, Gombert, Sebastian, Camus, Leon, Drachsler, Hendrik, Kroehne, Ulf |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Confidence Estimation in Automatic Short Answer Grading with LLMs
by: Cong, Longwei, et al.
Published: (2026)
by: Cong, Longwei, et al.
Published: (2026)
Enhancing LLM-Based Short Answer Grading with Retrieval-Augmented Generation
by: Chu, Yucheng, et al.
Published: (2025)
by: Chu, Yucheng, et al.
Published: (2025)
Grade Guard: A Smart System for Short Answer Automated Grading
by: Dadu, Niharika, et al.
Published: (2025)
by: Dadu, Niharika, et al.
Published: (2025)
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
by: Schmucker, Robin, et al.
Published: (2025)
by: Schmucker, Robin, et al.
Published: (2025)
Enhancing Multi-Domain Automatic Short Answer Grading through an Explainable Neuro-Symbolic Pipeline
by: Künnecke, Felix, et al.
Published: (2024)
by: Künnecke, Felix, et al.
Published: (2024)
"I understand why I got this grade": Automatic Short Answer Grading with Feedback
by: Aggarwal, Dishank, et al.
Published: (2024)
by: Aggarwal, Dishank, et al.
Published: (2024)
RIDE: Difficulty Evolving Perturbation with Item Response Theory for Mathematical Reasoning
by: Li, Xinyuan, et al.
Published: (2025)
by: Li, Xinyuan, et al.
Published: (2025)
Auditing LLM Benchmarks with Item Response Theory
by: Land, Sander, et al.
Published: (2026)
by: Land, Sander, et al.
Published: (2026)
ASAG2024: A Combined Benchmark for Short Answer Grading
by: Meyer, Gérôme, et al.
Published: (2024)
by: Meyer, Gérôme, et al.
Published: (2024)
Enhancing Security and Strengthening Defenses in Automated Short-Answer Grading Systems
by: Yarmohammadtoosky, Sahar, et al.
Published: (2025)
by: Yarmohammadtoosky, Sahar, et al.
Published: (2025)
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
by: Scarlatos, Alexander, et al.
Published: (2025)
by: Scarlatos, Alexander, et al.
Published: (2025)
JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory
by: Yao, Louie Hong, et al.
Published: (2025)
by: Yao, Louie Hong, et al.
Published: (2025)
CHiL(L)Grader: Calibrated Human-in-the-Loop Short-Answer Grading
by: Raikote, Pranav, et al.
Published: (2026)
by: Raikote, Pranav, et al.
Published: (2026)
From Flat to Structural: Enhancing Automated Short Answer Grading with GraphRAG
by: Chu, Yucheng, et al.
Published: (2026)
by: Chu, Yucheng, et al.
Published: (2026)
Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education
by: Henkel, Owen, et al.
Published: (2024)
by: Henkel, Owen, et al.
Published: (2024)
DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training
by: Tian, Xiaoyu, et al.
Published: (2025)
by: Tian, Xiaoyu, et al.
Published: (2025)
Automated Long Answer Grading with RiceChem Dataset
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
UnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory
by: Uebayashi, Shunki, et al.
Published: (2026)
by: Uebayashi, Shunki, et al.
Published: (2026)
Statistical Comparative Analysis of Semantic Similarities and Model Transferability Across Datasets for Short Answer Grading
by: Bonthu, Sridevi, et al.
Published: (2025)
by: Bonthu, Sridevi, et al.
Published: (2025)
Can LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel Dataset
by: Henkel, Owen, et al.
Published: (2023)
by: Henkel, Owen, et al.
Published: (2023)
Learning Compact Representations of LLM Abilities via Item Response Theory
by: Chen, Jianhao, et al.
Published: (2025)
by: Chen, Jianhao, et al.
Published: (2025)
Hey AI Can You Grade My Essay?: Automatic Essay Grading
by: Maliha, Maisha, et al.
Published: (2024)
by: Maliha, Maisha, et al.
Published: (2024)
Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations
by: Acquaye, Christabel, et al.
Published: (2026)
by: Acquaye, Christabel, et al.
Published: (2026)
Escaping binary gender roles: Gender diversity dynamics in a CSCL‐Escape game
by: Dana Kube, et al.
Published: (2024)
by: Dana Kube, et al.
Published: (2024)
Automatic Short‐Answer Grading in Sustainability Education: AI –Human Agreement
by: Emrah Emirtekin, et al.
Published: (2025)
by: Emrah Emirtekin, et al.
Published: (2025)
Focusing on Students, not Machines: Grounded Question Generation and Automated Answer Grading
by: Meyer, Gérôme, et al.
Published: (2025)
by: Meyer, Gérôme, et al.
Published: (2025)
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs
by: Furuhashi, Momoka, et al.
Published: (2025)
by: Furuhashi, Momoka, et al.
Published: (2025)
Question Difficulty Estimation for Large Language Models via Answer Plausibility Scoring
by: Mozafari, Jamshid, et al.
Published: (2026)
by: Mozafari, Jamshid, et al.
Published: (2026)
A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization
by: Chu, Yucheng, et al.
Published: (2024)
by: Chu, Yucheng, et al.
Published: (2024)
Prediction of Item Difficulty for Reading Comprehension Items by Creation of Annotated Item Repository
by: Kapoor, Radhika, et al.
Published: (2025)
by: Kapoor, Radhika, et al.
Published: (2025)
Explainable Automatic Grading with Neural Additive Models
by: Condor, Aubrey, et al.
Published: (2024)
by: Condor, Aubrey, et al.
Published: (2024)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
by: Dinh, Tu Anh, et al.
Published: (2024)
by: Dinh, Tu Anh, et al.
Published: (2024)
Japanese-English Sentence Translation Exercises Dataset for Automatic Grading
by: Miura, Naoki, et al.
Published: (2024)
by: Miura, Naoki, et al.
Published: (2024)
GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
by: Li, Xueyi, et al.
Published: (2026)
by: Li, Xueyi, et al.
Published: (2026)
LLM-based Automated Grading with Human-in-the-Loop
by: Chu, Yucheng, et al.
Published: (2025)
by: Chu, Yucheng, et al.
Published: (2025)
RocketEval: Efficient Automated LLM Evaluation via Grading Checklist
by: Wei, Tianjun, et al.
Published: (2025)
by: Wei, Tianjun, et al.
Published: (2025)
Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
by: Zhou, Hongli, et al.
Published: (2025)
by: Zhou, Hongli, et al.
Published: (2025)
Guarding the Privacy of Label-Only Access to Neural Network Classifiers via iDP Verification
by: Kabaha, Anan, et al.
Published: (2025)
by: Kabaha, Anan, et al.
Published: (2025)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Similar Items
-
Confidence Estimation in Automatic Short Answer Grading with LLMs
by: Cong, Longwei, et al.
Published: (2026) -
Enhancing LLM-Based Short Answer Grading with Retrieval-Augmented Generation
by: Chu, Yucheng, et al.
Published: (2025) -
Grade Guard: A Smart System for Short Answer Automated Grading
by: Dadu, Niharika, et al.
Published: (2025) -
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
by: Schmucker, Robin, et al.
Published: (2025) -
Enhancing Multi-Domain Automatic Short Answer Grading through an Explainable Neuro-Symbolic Pipeline
by: Künnecke, Felix, et al.
Published: (2024)