When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ferrer, Robinson, Turgut, Damla, Chen, Zhongzhou, Sonkar, Shashank |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
von: Do, Heejin, et al.
Veröffentlicht: (2026)
von: Do, Heejin, et al.
Veröffentlicht: (2026)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
von: Badawi, Abeer, et al.
Veröffentlicht: (2025)
von: Badawi, Abeer, et al.
Veröffentlicht: (2025)
MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education
von: Liu, Naiming, et al.
Veröffentlicht: (2024)
von: Liu, Naiming, et al.
Veröffentlicht: (2024)
LLM-based Cognitive Models of Students with Misconceptions
von: Sonkar, Shashank, et al.
Veröffentlicht: (2024)
von: Sonkar, Shashank, et al.
Veröffentlicht: (2024)
Atomic Learning Objectives Labeling: A High-Resolution Approach for Physics Education
von: Liu, Naiming, et al.
Veröffentlicht: (2024)
von: Liu, Naiming, et al.
Veröffentlicht: (2024)
FoundationalASSIST: An Educational Dataset for Foundational Knowledge Tracing and Pedagogical Grounding of LLMs
von: Worden, Eamon, et al.
Veröffentlicht: (2026)
von: Worden, Eamon, et al.
Veröffentlicht: (2026)
Scalable Generation and Validation of Isomorphic Physics Problems with GenAI
von: Liu, Naiming, et al.
Veröffentlicht: (2026)
von: Liu, Naiming, et al.
Veröffentlicht: (2026)
MetaCLASS: Metacognitive Coaching for Learning with Adaptive Self-regulation Support
von: Liu, Naiming, et al.
Veröffentlicht: (2026)
von: Liu, Naiming, et al.
Veröffentlicht: (2026)
Are Large Language Models Good Essay Graders?
von: Kundu, Anindita, et al.
Veröffentlicht: (2024)
von: Kundu, Anindita, et al.
Veröffentlicht: (2024)
Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring
von: Wu, Xuansheng, et al.
Veröffentlicht: (2024)
von: Wu, Xuansheng, et al.
Veröffentlicht: (2024)
MalruleLib: Large-Scale Executable Misconception Reasoning with Step Traces for Modeling Student Thinking in Mathematics
von: Chen, Xinghe, et al.
Veröffentlicht: (2026)
von: Chen, Xinghe, et al.
Veröffentlicht: (2026)
Many-Shot Regurgitation (MSR) Prompting
von: Sonkar, Shashank, et al.
Veröffentlicht: (2024)
von: Sonkar, Shashank, et al.
Veröffentlicht: (2024)
Toward LLM-Supported Automated Assessment of Critical Thinking Subskills
von: Peczuh, Marisa C., et al.
Veröffentlicht: (2025)
von: Peczuh, Marisa C., et al.
Veröffentlicht: (2025)
LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2026)
von: Akpinar, Nil-Jana, et al.
Veröffentlicht: (2026)
Can We Trust LLM Detectors?
von: Sandhan, Jivnesh, et al.
Veröffentlicht: (2026)
von: Sandhan, Jivnesh, et al.
Veröffentlicht: (2026)
Can LLM be a Personalized Judge?
von: Dong, Yijiang River, et al.
Veröffentlicht: (2024)
von: Dong, Yijiang River, et al.
Veröffentlicht: (2024)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
von: Kwon, Jea, et al.
Veröffentlicht: (2025)
von: Kwon, Jea, et al.
Veröffentlicht: (2025)
CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models
von: Liu, Naiming, et al.
Veröffentlicht: (2025)
von: Liu, Naiming, et al.
Veröffentlicht: (2025)
From Black-Box Confidence to Measurable Trust in Clinical AI: A Framework for Evidence, Supervision, and Staged Autonomy
von: Zabolotnii, Serhii, et al.
Veröffentlicht: (2026)
von: Zabolotnii, Serhii, et al.
Veröffentlicht: (2026)
LLM-REVal: Can We Trust LLM Reviewers Yet?
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Misconception Acquisition Dynamics in Large Language Models
von: Liu, Naiming, et al.
Veröffentlicht: (2026)
von: Liu, Naiming, et al.
Veröffentlicht: (2026)
Auto311: A Confidence-guided Automated System for Non-emergency Calls
von: Chen, Zirong, et al.
Veröffentlicht: (2023)
von: Chen, Zirong, et al.
Veröffentlicht: (2023)
When Neutral Summaries are not that Neutral: Quantifying Political Neutrality in LLM-Generated News Summaries
von: Vijay, Supriti, et al.
Veröffentlicht: (2024)
von: Vijay, Supriti, et al.
Veröffentlicht: (2024)
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
von: Zong, Qing, et al.
Veröffentlicht: (2025)
von: Zong, Qing, et al.
Veröffentlicht: (2025)
Student Data Paradox and Curious Case of Single Student-Tutor Model: Regressive Side Effects of Training LLMs for Personalized Learning
von: Sonkar, Shashank, et al.
Veröffentlicht: (2024)
von: Sonkar, Shashank, et al.
Veröffentlicht: (2024)
Do LLMs Make Mistakes Like Students? Exploring Natural Alignment between Language Models and Human Error Patterns
von: Liu, Naiming, et al.
Veröffentlicht: (2025)
von: Liu, Naiming, et al.
Veröffentlicht: (2025)
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
von: Hwang, Yerin, et al.
Veröffentlicht: (2025)
von: Hwang, Yerin, et al.
Veröffentlicht: (2025)
Can LLMs Reason About Trust?: A Pilot Study
von: Debnath, Anushka, et al.
Veröffentlicht: (2025)
von: Debnath, Anushka, et al.
Veröffentlicht: (2025)
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
von: Harvey, Emma, et al.
Veröffentlicht: (2025)
von: Harvey, Emma, et al.
Veröffentlicht: (2025)
Improving Clustering on Occupational Text Data through Dimensionality Reduction
von: García, Iago Xabier Vázquez, et al.
Veröffentlicht: (2025)
von: García, Iago Xabier Vázquez, et al.
Veröffentlicht: (2025)
On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral
von: Steenhuis, Quinten, et al.
Veröffentlicht: (2026)
von: Steenhuis, Quinten, et al.
Veröffentlicht: (2026)
Misclassification in Automated Content Analysis Causes Bias in Regression. Can We Fix It? Yes We Can!
von: TeBlunthuis, Nathan, et al.
Veröffentlicht: (2023)
von: TeBlunthuis, Nathan, et al.
Veröffentlicht: (2023)
Use Me Wisely: AI-Driven Assessment for LLM Prompting Skills Development
von: Ognibene, Dimitri, et al.
Veröffentlicht: (2025)
von: Ognibene, Dimitri, et al.
Veröffentlicht: (2025)
Safeguarding Decentralized Social Media: LLM Agents for Automating Community Rule Compliance
von: La Cava, Lucio, et al.
Veröffentlicht: (2024)
von: La Cava, Lucio, et al.
Veröffentlicht: (2024)
Automated Assessment of Students' Code Comprehension using LLMs
von: Oli, Priti, et al.
Veröffentlicht: (2023)
von: Oli, Priti, et al.
Veröffentlicht: (2023)
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
von: Wang, Leyao, et al.
Veröffentlicht: (2026)
von: Wang, Leyao, et al.
Veröffentlicht: (2026)
When to Trust LLMs: Aligning Confidence with Response Quality
von: Tao, Shuchang, et al.
Veröffentlicht: (2024)
von: Tao, Shuchang, et al.
Veröffentlicht: (2024)
Can AI Debias the News? LLM Interventions Improve Cross-Partisan Receptivity but LLMs Overestimate Their Own Effectiveness
von: Feroz, Faisal, et al.
Veröffentlicht: (2026)
von: Feroz, Faisal, et al.
Veröffentlicht: (2026)
Automated Long Answer Grading with RiceChem Dataset
von: Sonkar, Shashank, et al.
Veröffentlicht: (2024)
von: Sonkar, Shashank, et al.
Veröffentlicht: (2024)
Societal Alignment Frameworks Can Improve LLM Alignment
von: Stańczak, Karolina, et al.
Veröffentlicht: (2025)
von: Stańczak, Karolina, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
von: Do, Heejin, et al.
Veröffentlicht: (2026) -
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
von: Badawi, Abeer, et al.
Veröffentlicht: (2025) -
MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education
von: Liu, Naiming, et al.
Veröffentlicht: (2024) -
LLM-based Cognitive Models of Students with Misconceptions
von: Sonkar, Shashank, et al.
Veröffentlicht: (2024) -
Atomic Learning Objectives Labeling: A High-Resolution Approach for Physics Education
von: Liu, Naiming, et al.
Veröffentlicht: (2024)