When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
Fuente:
arXiv
Guardado en:
| Autores principales: | Ferrer, Robinson, Turgut, Damla, Chen, Zhongzhou, Sonkar, Shashank |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
por: Do, Heejin, et al.
Publicado: (2026)
por: Do, Heejin, et al.
Publicado: (2026)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
por: Badawi, Abeer, et al.
Publicado: (2025)
por: Badawi, Abeer, et al.
Publicado: (2025)
MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education
por: Liu, Naiming, et al.
Publicado: (2024)
por: Liu, Naiming, et al.
Publicado: (2024)
LLM-based Cognitive Models of Students with Misconceptions
por: Sonkar, Shashank, et al.
Publicado: (2024)
por: Sonkar, Shashank, et al.
Publicado: (2024)
Atomic Learning Objectives Labeling: A High-Resolution Approach for Physics Education
por: Liu, Naiming, et al.
Publicado: (2024)
por: Liu, Naiming, et al.
Publicado: (2024)
FoundationalASSIST: An Educational Dataset for Foundational Knowledge Tracing and Pedagogical Grounding of LLMs
por: Worden, Eamon, et al.
Publicado: (2026)
por: Worden, Eamon, et al.
Publicado: (2026)
Scalable Generation and Validation of Isomorphic Physics Problems with GenAI
por: Liu, Naiming, et al.
Publicado: (2026)
por: Liu, Naiming, et al.
Publicado: (2026)
MetaCLASS: Metacognitive Coaching for Learning with Adaptive Self-regulation Support
por: Liu, Naiming, et al.
Publicado: (2026)
por: Liu, Naiming, et al.
Publicado: (2026)
Are Large Language Models Good Essay Graders?
por: Kundu, Anindita, et al.
Publicado: (2024)
por: Kundu, Anindita, et al.
Publicado: (2024)
Unveiling Scoring Processes: Dissecting the Differences between LLMs and Human Graders in Automatic Scoring
por: Wu, Xuansheng, et al.
Publicado: (2024)
por: Wu, Xuansheng, et al.
Publicado: (2024)
MalruleLib: Large-Scale Executable Misconception Reasoning with Step Traces for Modeling Student Thinking in Mathematics
por: Chen, Xinghe, et al.
Publicado: (2026)
por: Chen, Xinghe, et al.
Publicado: (2026)
Many-Shot Regurgitation (MSR) Prompting
por: Sonkar, Shashank, et al.
Publicado: (2024)
por: Sonkar, Shashank, et al.
Publicado: (2024)
Toward LLM-Supported Automated Assessment of Critical Thinking Subskills
por: Peczuh, Marisa C., et al.
Publicado: (2025)
por: Peczuh, Marisa C., et al.
Publicado: (2025)
LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
por: Akpinar, Nil-Jana, et al.
Publicado: (2026)
por: Akpinar, Nil-Jana, et al.
Publicado: (2026)
Can We Trust LLM Detectors?
por: Sandhan, Jivnesh, et al.
Publicado: (2026)
por: Sandhan, Jivnesh, et al.
Publicado: (2026)
Can LLM be a Personalized Judge?
por: Dong, Yijiang River, et al.
Publicado: (2024)
por: Dong, Yijiang River, et al.
Publicado: (2024)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
por: Kwon, Jea, et al.
Publicado: (2025)
por: Kwon, Jea, et al.
Publicado: (2025)
CLEAR-3K: Assessing Causal Explanatory Capabilities in Language Models
por: Liu, Naiming, et al.
Publicado: (2025)
por: Liu, Naiming, et al.
Publicado: (2025)
From Black-Box Confidence to Measurable Trust in Clinical AI: A Framework for Evidence, Supervision, and Staged Autonomy
por: Zabolotnii, Serhii, et al.
Publicado: (2026)
por: Zabolotnii, Serhii, et al.
Publicado: (2026)
LLM-REVal: Can We Trust LLM Reviewers Yet?
por: Li, Rui, et al.
Publicado: (2025)
por: Li, Rui, et al.
Publicado: (2025)
Misconception Acquisition Dynamics in Large Language Models
por: Liu, Naiming, et al.
Publicado: (2026)
por: Liu, Naiming, et al.
Publicado: (2026)
Auto311: A Confidence-guided Automated System for Non-emergency Calls
por: Chen, Zirong, et al.
Publicado: (2023)
por: Chen, Zirong, et al.
Publicado: (2023)
When Neutral Summaries are not that Neutral: Quantifying Political Neutrality in LLM-Generated News Summaries
por: Vijay, Supriti, et al.
Publicado: (2024)
por: Vijay, Supriti, et al.
Publicado: (2024)
CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?
por: Zong, Qing, et al.
Publicado: (2025)
por: Zong, Qing, et al.
Publicado: (2025)
Student Data Paradox and Curious Case of Single Student-Tutor Model: Regressive Side Effects of Training LLMs for Personalized Learning
por: Sonkar, Shashank, et al.
Publicado: (2024)
por: Sonkar, Shashank, et al.
Publicado: (2024)
Do LLMs Make Mistakes Like Students? Exploring Natural Alignment between Language Models and Human Error Patterns
por: Liu, Naiming, et al.
Publicado: (2025)
por: Liu, Naiming, et al.
Publicado: (2025)
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
por: Hwang, Yerin, et al.
Publicado: (2025)
por: Hwang, Yerin, et al.
Publicado: (2025)
Can LLMs Reason About Trust?: A Pilot Study
por: Debnath, Anushka, et al.
Publicado: (2025)
por: Debnath, Anushka, et al.
Publicado: (2025)
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
por: Harvey, Emma, et al.
Publicado: (2025)
por: Harvey, Emma, et al.
Publicado: (2025)
Improving Clustering on Occupational Text Data through Dimensionality Reduction
por: García, Iago Xabier Vázquez, et al.
Publicado: (2025)
por: García, Iago Xabier Vázquez, et al.
Publicado: (2025)
On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral
por: Steenhuis, Quinten, et al.
Publicado: (2026)
por: Steenhuis, Quinten, et al.
Publicado: (2026)
Misclassification in Automated Content Analysis Causes Bias in Regression. Can We Fix It? Yes We Can!
por: TeBlunthuis, Nathan, et al.
Publicado: (2023)
por: TeBlunthuis, Nathan, et al.
Publicado: (2023)
Use Me Wisely: AI-Driven Assessment for LLM Prompting Skills Development
por: Ognibene, Dimitri, et al.
Publicado: (2025)
por: Ognibene, Dimitri, et al.
Publicado: (2025)
Safeguarding Decentralized Social Media: LLM Agents for Automating Community Rule Compliance
por: La Cava, Lucio, et al.
Publicado: (2024)
por: La Cava, Lucio, et al.
Publicado: (2024)
Automated Assessment of Students' Code Comprehension using LLMs
por: Oli, Priti, et al.
Publicado: (2023)
por: Oli, Priti, et al.
Publicado: (2023)
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
por: Wang, Leyao, et al.
Publicado: (2026)
por: Wang, Leyao, et al.
Publicado: (2026)
When to Trust LLMs: Aligning Confidence with Response Quality
por: Tao, Shuchang, et al.
Publicado: (2024)
por: Tao, Shuchang, et al.
Publicado: (2024)
Can AI Debias the News? LLM Interventions Improve Cross-Partisan Receptivity but LLMs Overestimate Their Own Effectiveness
por: Feroz, Faisal, et al.
Publicado: (2026)
por: Feroz, Faisal, et al.
Publicado: (2026)
Automated Long Answer Grading with RiceChem Dataset
por: Sonkar, Shashank, et al.
Publicado: (2024)
por: Sonkar, Shashank, et al.
Publicado: (2024)
Societal Alignment Frameworks Can Improve LLM Alignment
por: Stańczak, Karolina, et al.
Publicado: (2025)
por: Stańczak, Karolina, et al.
Publicado: (2025)
Ejemplares similares
-
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
por: Do, Heejin, et al.
Publicado: (2026) -
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
por: Badawi, Abeer, et al.
Publicado: (2025) -
MalAlgoQA: Pedagogical Evaluation of Counterfactual Reasoning in Large Language Models and Implications for AI in Education
por: Liu, Naiming, et al.
Publicado: (2024) -
LLM-based Cognitive Models of Students with Misconceptions
por: Sonkar, Shashank, et al.
Publicado: (2024) -
Atomic Learning Objectives Labeling: A High-Resolution Approach for Physics Education
por: Liu, Naiming, et al.
Publicado: (2024)