How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Hang, Yang, Kaiqi, Long, Xianxuan, Filippov, Fedor, Chu, Yucheng, Copur-Gencturk, Yasemin, He, Peng, Miller, Cory, Shin, Namsoo, Krajcik, Joseph, Liu, Hui, Tang, Jiliang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Confusion-Aware Rubric Optimization for LLM-based Automated Grading
di: Chu, Yucheng, et al.
Pubblicazione: (2026)
di: Chu, Yucheng, et al.
Pubblicazione: (2026)
Optimizing In-Context Demonstrations for LLM-based Automated Grading
di: Chu, Yucheng, et al.
Pubblicazione: (2026)
di: Chu, Yucheng, et al.
Pubblicazione: (2026)
From Flat to Structural: Enhancing Automated Short Answer Grading with GraphRAG
di: Chu, Yucheng, et al.
Pubblicazione: (2026)
di: Chu, Yucheng, et al.
Pubblicazione: (2026)
LLM-based Automated Grading with Human-in-the-Loop
di: Chu, Yucheng, et al.
Pubblicazione: (2025)
di: Chu, Yucheng, et al.
Pubblicazione: (2025)
A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization
di: Chu, Yucheng, et al.
Pubblicazione: (2024)
di: Chu, Yucheng, et al.
Pubblicazione: (2024)
An Artificial Intelligence‐Enhanced Assessment Framework for Analyzing Middle School Science Students’ Written Responses
di: Namsoo Shin, et al.
Pubblicazione: (2026)
di: Namsoo Shin, et al.
Pubblicazione: (2026)
A LLM-Driven Multi-Agent Systems for Professional Development of Mathematics Teachers
di: Yang, Kaiqi, et al.
Pubblicazione: (2025)
di: Yang, Kaiqi, et al.
Pubblicazione: (2025)
Content Knowledge Identification with Multi-Agent Large Language Models (LLMs)
di: Yang, Kaiqi, et al.
Pubblicazione: (2024)
di: Yang, Kaiqi, et al.
Pubblicazione: (2024)
Enhancing LLM-Based Short Answer Grading with Retrieval-Augmented Generation
di: Chu, Yucheng, et al.
Pubblicazione: (2025)
di: Chu, Yucheng, et al.
Pubblicazione: (2025)
Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
di: Li, Hang, et al.
Pubblicazione: (2025)
di: Li, Hang, et al.
Pubblicazione: (2025)
Assessing the Nature of Learners' Science Content Understandings as a Result of Utilizing On-Line Resources.
di: Hoffman, Joseph L., et al.
Pubblicazione: (1999)
di: Hoffman, Joseph L., et al.
Pubblicazione: (1999)
Iterative LLM-Based Generation and Refinement of Distracting Conditions in Math Word Problems
di: Yang, Kaiqi, et al.
Pubblicazione: (2025)
di: Yang, Kaiqi, et al.
Pubblicazione: (2025)
Factors predicting teachers' implementation of inquiry‐based teaching practices: Analysis of South African TIMSS 2019 data from an ecological perspective
di: Ayodele Abosede Ogegbo, et al.
Pubblicazione: (2024)
di: Ayodele Abosede Ogegbo, et al.
Pubblicazione: (2024)
Effectiveness of Ozone Therapy in Botulinum Toxin–Induced Ptosis: Two Case Reports
di: Selda Yıldırım Gençtürk, et al.
Pubblicazione: (2026)
di: Selda Yıldırım Gençtürk, et al.
Pubblicazione: (2026)
Grading Standards in Education Departments at Universities
di: Cory Koedel
Pubblicazione: (2011)
di: Cory Koedel
Pubblicazione: (2011)
Selected Traits of Hatched and Unhatched Eggs and Growth Performance of Yellow Japanese Quails
di: G Copur Akpinar
Pubblicazione: (2017)
di: G Copur Akpinar
Pubblicazione: (2017)
FairlyUncertain: A Comprehensive Benchmark of Uncertainty in Algorithmic Fairness
di: Rosenblatt, Lucas, et al.
Pubblicazione: (2024)
di: Rosenblatt, Lucas, et al.
Pubblicazione: (2024)
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
di: Chandu, Khyathi Raghavi, et al.
Pubblicazione: (2024)
di: Chandu, Khyathi Raghavi, et al.
Pubblicazione: (2024)
Beyond Partisan Leaning: A Comparative Analysis of Political Bias in Large Language Models
di: Peng, Tai-Quan, et al.
Pubblicazione: (2024)
di: Peng, Tai-Quan, et al.
Pubblicazione: (2024)
Are Large Language Models (LLMs) Good Social Predictors?
di: Yang, Kaiqi, et al.
Pubblicazione: (2024)
di: Yang, Kaiqi, et al.
Pubblicazione: (2024)
GraphGhost: Tracing Structures Behind Large Language Models
di: Dai, Xinnan, et al.
Pubblicazione: (2025)
di: Dai, Xinnan, et al.
Pubblicazione: (2025)
Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem Solutions
di: Li, Hang, et al.
Pubblicazione: (2024)
di: Li, Hang, et al.
Pubblicazione: (2024)
Learning Progression-Guided AI Evaluation of Scientific Models To Support Diverse Multi-Modal Understanding in NGSS Classroom
di: Kaldaras, Leonora, et al.
Pubblicazione: (2025)
di: Kaldaras, Leonora, et al.
Pubblicazione: (2025)
Reasoning by Exploration: A Unified Approach to Retrieval and Generation over Graphs
di: Han, Haoyu, et al.
Pubblicazione: (2025)
di: Han, Haoyu, et al.
Pubblicazione: (2025)
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
di: Fandina, Ora Nova, et al.
Pubblicazione: (2024)
di: Fandina, Ora Nova, et al.
Pubblicazione: (2024)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
di: Dinh, Tu Anh, et al.
Pubblicazione: (2024)
di: Dinh, Tu Anh, et al.
Pubblicazione: (2024)
Benchmarking Usage Statistics in Collection Management Decisions for Serials
di: Tucker, Cory
Pubblicazione: (2009)
di: Tucker, Cory
Pubblicazione: (2009)
Graded strength of comparative illusions is explained by Bayesian inference
di: Zhang, Yuhan, et al.
Pubblicazione: (2025)
di: Zhang, Yuhan, et al.
Pubblicazione: (2025)
Bridging Local and Federated Data Normalization in Federated Learning: A Privacy-Preserving Approach
di: Coşğun, Melih, et al.
Pubblicazione: (2025)
di: Coşğun, Melih, et al.
Pubblicazione: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
di: Yang, Langqi, et al.
Pubblicazione: (2025)
di: Yang, Langqi, et al.
Pubblicazione: (2025)
The relationship between ethnocentrism and xenophobia level and predictors: A descriptive and correlational study of nurses working in two cities where refugees live intensively in Turkey
di: İpek Köse Tosunöz, et al.
Pubblicazione: (2024)
di: İpek Köse Tosunöz, et al.
Pubblicazione: (2024)
Paper WR v1.0: Geometric Optics Without Metric — Write Resistance and Photon Paths in Ontological Resolution Theory
di: Kapitanov, Fedor
Pubblicazione: (2026)
di: Kapitanov, Fedor
Pubblicazione: (2026)
Perfecting Depth: Uncertainty-Aware Enhancement of Metric Depth
di: Jun, Jinyoung, et al.
Pubblicazione: (2025)
di: Jun, Jinyoung, et al.
Pubblicazione: (2025)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
di: Ryan, Michael J., et al.
Pubblicazione: (2025)
di: Ryan, Michael J., et al.
Pubblicazione: (2025)
Transformer Circuit Faithfulness Metrics are not Robust
di: Miller, Joseph, et al.
Pubblicazione: (2024)
di: Miller, Joseph, et al.
Pubblicazione: (2024)
Automatic structures and the problem of natural well-orderings
di: Beklemishev, Lev D., et al.
Pubblicazione: (2024)
di: Beklemishev, Lev D., et al.
Pubblicazione: (2024)
Case Report: An Unusual Mimicker of Osteoarthritis
di: Sidar Çöpür, et al.
Pubblicazione: (2025)
di: Sidar Çöpür, et al.
Pubblicazione: (2025)
Paper V: The Execution Mechanism — How the Unique Executable Program Unfolds on the FCC Lattice
di: Kapitanov, Fedor
Pubblicazione: (2026)
di: Kapitanov, Fedor
Pubblicazione: (2026)
Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks
di: Longjohn, Rachel, et al.
Pubblicazione: (2025)
di: Longjohn, Rachel, et al.
Pubblicazione: (2025)
Automatic deep learning segmentation of the hippocampus on high‐resolution diffusion magnetic resonance imaging and its application to the healthy lifespan
di: Dylan Miller, et al.
Pubblicazione: (2024)
di: Dylan Miller, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Confusion-Aware Rubric Optimization for LLM-based Automated Grading
di: Chu, Yucheng, et al.
Pubblicazione: (2026) -
Optimizing In-Context Demonstrations for LLM-based Automated Grading
di: Chu, Yucheng, et al.
Pubblicazione: (2026) -
From Flat to Structural: Enhancing Automated Short Answer Grading with GraphRAG
di: Chu, Yucheng, et al.
Pubblicazione: (2026) -
LLM-based Automated Grading with Human-in-the-Loop
di: Chu, Yucheng, et al.
Pubblicazione: (2025) -
A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization
di: Chu, Yucheng, et al.
Pubblicazione: (2024)