Confusion-Aware Rubric Optimization for LLM-based Automated Grading
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chu, Yucheng, Li, Hang, Yang, Kaiqi, Copur-Gencturk, Yasemin, Krajcik, Joseph, Shin, Namsoo, Tang, Jiliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimizing In-Context Demonstrations for LLM-based Automated Grading
von: Chu, Yucheng, et al.
Veröffentlicht: (2026)
von: Chu, Yucheng, et al.
Veröffentlicht: (2026)
From Flat to Structural: Enhancing Automated Short Answer Grading with GraphRAG
von: Chu, Yucheng, et al.
Veröffentlicht: (2026)
von: Chu, Yucheng, et al.
Veröffentlicht: (2026)
LLM-based Automated Grading with Human-in-the-Loop
von: Chu, Yucheng, et al.
Veröffentlicht: (2025)
von: Chu, Yucheng, et al.
Veröffentlicht: (2025)
A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization
von: Chu, Yucheng, et al.
Veröffentlicht: (2024)
von: Chu, Yucheng, et al.
Veröffentlicht: (2024)
How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment
von: Li, Hang, et al.
Veröffentlicht: (2026)
von: Li, Hang, et al.
Veröffentlicht: (2026)
Content Knowledge Identification with Multi-Agent Large Language Models (LLMs)
von: Yang, Kaiqi, et al.
Veröffentlicht: (2024)
von: Yang, Kaiqi, et al.
Veröffentlicht: (2024)
Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
von: Li, Hang, et al.
Veröffentlicht: (2025)
von: Li, Hang, et al.
Veröffentlicht: (2025)
Enhancing LLM-Based Short Answer Grading with Retrieval-Augmented Generation
von: Chu, Yucheng, et al.
Veröffentlicht: (2025)
von: Chu, Yucheng, et al.
Veröffentlicht: (2025)
A LLM-Driven Multi-Agent Systems for Professional Development of Mathematics Teachers
von: Yang, Kaiqi, et al.
Veröffentlicht: (2025)
von: Yang, Kaiqi, et al.
Veröffentlicht: (2025)
Autorubric: Unifying Rubric-based LLM Evaluation
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
An Artificial Intelligence‐Enhanced Assessment Framework for Analyzing Middle School Science Students’ Written Responses
von: Namsoo Shin, et al.
Veröffentlicht: (2026)
von: Namsoo Shin, et al.
Veröffentlicht: (2026)
AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning
von: Ding, Liang
Veröffentlicht: (2026)
von: Ding, Liang
Veröffentlicht: (2026)
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
von: Chen, Yinzhu, et al.
Veröffentlicht: (2026)
von: Chen, Yinzhu, et al.
Veröffentlicht: (2026)
REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading
von: Zhao, Chengshuai, et al.
Veröffentlicht: (2026)
von: Zhao, Chengshuai, et al.
Veröffentlicht: (2026)
Knowledge Tagging System on Math Questions via LLMs with Flexible Demonstration Retriever
von: Li, Hang, et al.
Veröffentlicht: (2024)
von: Li, Hang, et al.
Veröffentlicht: (2024)
Designing Reliable LLM-Assisted Rubric Scoring for Constructed Responses: Evidence from Physics Exams
von: Tang, Xiuxiu, et al.
Veröffentlicht: (2026)
von: Tang, Xiuxiu, et al.
Veröffentlicht: (2026)
Optimsyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation
von: Fan, Zhiting, et al.
Veröffentlicht: (2026)
von: Fan, Zhiting, et al.
Veröffentlicht: (2026)
When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025)
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025)
LASER: An LLM-based ASR Scoring and Evaluation Rubric
von: Parulekar, Amruta, et al.
Veröffentlicht: (2025)
von: Parulekar, Amruta, et al.
Veröffentlicht: (2025)
Kardia-R1: Unleashing LLMs to Reason toward Understanding and Empathy for Emotional Support via Rubric-as-Judge Reinforcement Learning
von: Yuan, Jiahao, et al.
Veröffentlicht: (2025)
von: Yuan, Jiahao, et al.
Veröffentlicht: (2025)
LLM Essay Scoring Under Holistic and Analytic Rubrics: Prompt Effects and Bias
von: Kucia, Filip J., et al.
Veröffentlicht: (2026)
von: Kucia, Filip J., et al.
Veröffentlicht: (2026)
DREsS: Dataset for Rubric-based Essay Scoring on EFL Writing
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
von: Yoo, Haneul, et al.
Veröffentlicht: (2024)
TALEC: Teach Your LLM to Evaluate in Specific Domain with In-house Criteria by Criteria Division and Zero-shot Plus Few-shot
von: Zhang, Kaiqi, et al.
Veröffentlicht: (2024)
von: Zhang, Kaiqi, et al.
Veröffentlicht: (2024)
Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks
von: Xu, Tianze, et al.
Veröffentlicht: (2026)
von: Xu, Tianze, et al.
Veröffentlicht: (2026)
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026)
von: Ding, Ruomeng, et al.
Veröffentlicht: (2026)
Listwise Preference Optimization with Element-wise Confusions for Aspect Sentiment Quad Prediction
von: Lai, Wenna, et al.
Veröffentlicht: (2025)
von: Lai, Wenna, et al.
Veröffentlicht: (2025)
Rubric-Guided Process Reward for Stepwise Model Routing
von: Ye, Shenghao, et al.
Veröffentlicht: (2026)
von: Ye, Shenghao, et al.
Veröffentlicht: (2026)
MAC-Tuning: LLM Multi-Compositional Problem Reasoning with Enhanced Knowledge Boundary Awareness
von: Huang, Junsheng, et al.
Veröffentlicht: (2025)
von: Huang, Junsheng, et al.
Veröffentlicht: (2025)
What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance
von: Watson, William, et al.
Veröffentlicht: (2026)
von: Watson, William, et al.
Veröffentlicht: (2026)
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
von: Sharma, Manasi, et al.
Veröffentlicht: (2025)
von: Sharma, Manasi, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Rubric Anchors
von: Huang, Zenan, et al.
Veröffentlicht: (2025)
von: Huang, Zenan, et al.
Veröffentlicht: (2025)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
A Theoretical Understanding of Chain-of-Thought: Coherent Reasoning and Error-Aware Demonstration
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
LLMs can be easily Confused by Instructional Distractions
von: Hwang, Yerin, et al.
Veröffentlicht: (2025)
von: Hwang, Yerin, et al.
Veröffentlicht: (2025)
An Efficient Rubric-based Generative Verifier for Search-Augmented LLMs
von: Ma, Linyue, et al.
Veröffentlicht: (2025)
von: Ma, Linyue, et al.
Veröffentlicht: (2025)
AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning
von: Wu, Peilin, et al.
Veröffentlicht: (2026)
von: Wu, Peilin, et al.
Veröffentlicht: (2026)
Automated Survey Collection with LLM-based Conversational Agents
von: Kaiyrbekov, Kurmanbek, et al.
Veröffentlicht: (2025)
von: Kaiyrbekov, Kurmanbek, et al.
Veröffentlicht: (2025)
Configurable Preference Tuning with Rubric-Guided Synthetic Data
von: Gallego, Víctor
Veröffentlicht: (2025)
von: Gallego, Víctor
Veröffentlicht: (2025)
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation
von: Liu, Xue, et al.
Veröffentlicht: (2026)
von: Liu, Xue, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Optimizing In-Context Demonstrations for LLM-based Automated Grading
von: Chu, Yucheng, et al.
Veröffentlicht: (2026) -
From Flat to Structural: Enhancing Automated Short Answer Grading with GraphRAG
von: Chu, Yucheng, et al.
Veröffentlicht: (2026) -
LLM-based Automated Grading with Human-in-the-Loop
von: Chu, Yucheng, et al.
Veröffentlicht: (2025) -
A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization
von: Chu, Yucheng, et al.
Veröffentlicht: (2024) -
How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment
von: Li, Hang, et al.
Veröffentlicht: (2026)