Saved in:
| Main Authors: | Cong, Longwei, Hahn, Sonja, Gombert, Sebastian, Camus, Leon, Drachsler, Hendrik, Kroehne, Ulf |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.00200 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory
by: Cong, Longwei, et al.
Published: (2026)
by: Cong, Longwei, et al.
Published: (2026)
Escaping binary gender roles: Gender diversity dynamics in a CSCL‐Escape game
by: Dana Kube, et al.
Published: (2024)
by: Dana Kube, et al.
Published: (2024)
"I understand why I got this grade": Automatic Short Answer Grading with Feedback
by: Aggarwal, Dishank, et al.
Published: (2024)
by: Aggarwal, Dishank, et al.
Published: (2024)
Enhancing Multi-Domain Automatic Short Answer Grading through an Explainable Neuro-Symbolic Pipeline
by: Künnecke, Felix, et al.
Published: (2024)
by: Künnecke, Felix, et al.
Published: (2024)
ADVICE: Answer-Dependent Verbalized Confidence Estimation
by: Seo, Ki Jung, et al.
Published: (2025)
by: Seo, Ki Jung, et al.
Published: (2025)
Automatic Feedback Generation for Short Answer Questions using Answer Diagnostic Graphs
by: Furuhashi, Momoka, et al.
Published: (2025)
by: Furuhashi, Momoka, et al.
Published: (2025)
Grade Guard: A Smart System for Short Answer Automated Grading
by: Dadu, Niharika, et al.
Published: (2025)
by: Dadu, Niharika, et al.
Published: (2025)
Enhancing LLM-Based Short Answer Grading with Retrieval-Augmented Generation
by: Chu, Yucheng, et al.
Published: (2025)
by: Chu, Yucheng, et al.
Published: (2025)
Can LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel Dataset
by: Henkel, Owen, et al.
Published: (2023)
by: Henkel, Owen, et al.
Published: (2023)
CHiL(L)Grader: Calibrated Human-in-the-Loop Short-Answer Grading
by: Raikote, Pranav, et al.
Published: (2026)
by: Raikote, Pranav, et al.
Published: (2026)
ASAG2024: A Combined Benchmark for Short Answer Grading
by: Meyer, Gérôme, et al.
Published: (2024)
by: Meyer, Gérôme, et al.
Published: (2024)
Enhancing Security and Strengthening Defenses in Automated Short-Answer Grading Systems
by: Yarmohammadtoosky, Sahar, et al.
Published: (2025)
by: Yarmohammadtoosky, Sahar, et al.
Published: (2025)
From Flat to Structural: Enhancing Automated Short Answer Grading with GraphRAG
by: Chu, Yucheng, et al.
Published: (2026)
by: Chu, Yucheng, et al.
Published: (2026)
Learning Analytics in Higher Education -- Exploring Students and Teachers Expectations in Germany
by: Fritz, Birthe, et al.
Published: (2024)
by: Fritz, Birthe, et al.
Published: (2024)
Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education
by: Henkel, Owen, et al.
Published: (2024)
by: Henkel, Owen, et al.
Published: (2024)
Confidence Estimation for LLMs in Multi-turn Interactions
by: Zhang, Caiqi, et al.
Published: (2026)
by: Zhang, Caiqi, et al.
Published: (2026)
I don't have time! But keep me in the loop: Co‐designing requirements for a learning analytics cockpit with teachers
by: Onur Karademir, et al.
Published: (2024)
by: Onur Karademir, et al.
Published: (2024)
Statistical Comparative Analysis of Semantic Similarities and Model Transferability Across Datasets for Short Answer Grading
by: Bonthu, Sridevi, et al.
Published: (2025)
by: Bonthu, Sridevi, et al.
Published: (2025)
Beyond Scores: A Modular RAG-Based System for Automatic Short Answer Scoring with Feedback
by: Fateen, Menna, et al.
Published: (2024)
by: Fateen, Menna, et al.
Published: (2024)
Automated Long Answer Grading with RiceChem Dataset
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
by: Wang, Yuhan, et al.
Published: (2026)
by: Wang, Yuhan, et al.
Published: (2026)
Confidence Estimation for Automatic Detection of Depression and Alzheimer's Disease Based on Clinical Interviews
by: Wu, Wen, et al.
Published: (2024)
by: Wu, Wen, et al.
Published: (2024)
Mini-Batch Robustness Verification of Deep Neural Networks
by: Tzour-Shaday, Saar, et al.
Published: (2025)
by: Tzour-Shaday, Saar, et al.
Published: (2025)
The Curious Case of Factual (Mis)Alignment between LLMs' Short- and Long-Form Answers
by: Islam, Saad Obaid ul, et al.
Published: (2025)
by: Islam, Saad Obaid ul, et al.
Published: (2025)
Factual Confidence of LLMs: on Reliability and Robustness of Current Estimators
by: Mahaut, Matéo, et al.
Published: (2024)
by: Mahaut, Matéo, et al.
Published: (2024)
Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?
by: Lewis-Lim, Samuel, et al.
Published: (2025)
by: Lewis-Lim, Samuel, et al.
Published: (2025)
Automatic Question-Answer Generation for Long-Tail Knowledge
by: Kumar, Rohan, et al.
Published: (2024)
by: Kumar, Rohan, et al.
Published: (2024)
Automatic Short‐Answer Grading in Sustainability Education: AI –Human Agreement
by: Emrah Emirtekin, et al.
Published: (2025)
by: Emrah Emirtekin, et al.
Published: (2025)
Concept Map Assessment Through Structure Classification
by: Vossen, Laís P. V., et al.
Published: (2025)
by: Vossen, Laís P. V., et al.
Published: (2025)
Focusing on Students, not Machines: Grounded Question Generation and Automated Answer Grading
by: Meyer, Gérôme, et al.
Published: (2025)
by: Meyer, Gérôme, et al.
Published: (2025)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
Acting Flatterers via LLMs Sycophancy: Combating Clickbait with LLMs Opposing-Stance Reasoning
by: Zhang, Chaowei, et al.
Published: (2026)
by: Zhang, Chaowei, et al.
Published: (2026)
On Verbalized Confidence Scores for LLMs
by: Yang, Daniel, et al.
Published: (2024)
by: Yang, Daniel, et al.
Published: (2024)
Computational methods for Dynamic Answer Set Programming
by: Hahn, Susana
Published: (2025)
by: Hahn, Susana
Published: (2025)
Towards LLM-based Autograding for Short Textual Answers
by: Schneider, Johannes, et al.
Published: (2023)
by: Schneider, Johannes, et al.
Published: (2023)
LLMs Provide Unstable Answers to Legal Questions
by: Blair-Stanek, Andrew, et al.
Published: (2025)
by: Blair-Stanek, Andrew, et al.
Published: (2025)
Recursive Think-Answer Process for LLMs and VLMs
by: Lee, Byung-Kwan, et al.
Published: (2026)
by: Lee, Byung-Kwan, et al.
Published: (2026)
InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States
by: Beigi, Mohammad, et al.
Published: (2024)
by: Beigi, Mohammad, et al.
Published: (2024)
Reducing the Cost: Cross-Prompt Pre-Finetuning for Short Answer Scoring
by: Funayama, Hiroaki, et al.
Published: (2024)
by: Funayama, Hiroaki, et al.
Published: (2024)
Hallucination-Free Automatic Question & Answer Generation for Intuitive Learning
by: Wang, Nicholas X., et al.
Published: (2026)
by: Wang, Nicholas X., et al.
Published: (2026)
Similar Items
-
Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory
by: Cong, Longwei, et al.
Published: (2026) -
Escaping binary gender roles: Gender diversity dynamics in a CSCL‐Escape game
by: Dana Kube, et al.
Published: (2024) -
"I understand why I got this grade": Automatic Short Answer Grading with Feedback
by: Aggarwal, Dishank, et al.
Published: (2024) -
Enhancing Multi-Domain Automatic Short Answer Grading through an Explainable Neuro-Symbolic Pipeline
by: Künnecke, Felix, et al.
Published: (2024) -
ADVICE: Answer-Dependent Verbalized Confidence Estimation
by: Seo, Ki Jung, et al.
Published: (2025)