A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization
Fuente:
arXiv
Guardado en:
| Autores principales: | Chu, Yucheng, Li, Hang, Yang, Kaiqi, Shomer, Harry, Liu, Hui, Copur-Gencturk, Yasemin, Tang, Jiliang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLM-based Automated Grading with Human-in-the-Loop
por: Chu, Yucheng, et al.
Publicado: (2025)
por: Chu, Yucheng, et al.
Publicado: (2025)
Optimizing In-Context Demonstrations for LLM-based Automated Grading
por: Chu, Yucheng, et al.
Publicado: (2026)
por: Chu, Yucheng, et al.
Publicado: (2026)
Confusion-Aware Rubric Optimization for LLM-based Automated Grading
por: Chu, Yucheng, et al.
Publicado: (2026)
por: Chu, Yucheng, et al.
Publicado: (2026)
From Flat to Structural: Enhancing Automated Short Answer Grading with GraphRAG
por: Chu, Yucheng, et al.
Publicado: (2026)
por: Chu, Yucheng, et al.
Publicado: (2026)
How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment
por: Li, Hang, et al.
Publicado: (2026)
por: Li, Hang, et al.
Publicado: (2026)
A LLM-Driven Multi-Agent Systems for Professional Development of Mathematics Teachers
por: Yang, Kaiqi, et al.
Publicado: (2025)
por: Yang, Kaiqi, et al.
Publicado: (2025)
Content Knowledge Identification with Multi-Agent Large Language Models (LLMs)
por: Yang, Kaiqi, et al.
Publicado: (2024)
por: Yang, Kaiqi, et al.
Publicado: (2024)
Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
por: Li, Hang, et al.
Publicado: (2025)
por: Li, Hang, et al.
Publicado: (2025)
Enhancing LLM-Based Short Answer Grading with Retrieval-Augmented Generation
por: Chu, Yucheng, et al.
Publicado: (2025)
por: Chu, Yucheng, et al.
Publicado: (2025)
Towards Understanding Link Predictor Generalizability Under Distribution Shifts
por: Revolinsky, Jay, et al.
Publicado: (2024)
por: Revolinsky, Jay, et al.
Publicado: (2024)
Towards Better Benchmark Datasets for Inductive Knowledge Graph Completion
por: Shomer, Harry, et al.
Publicado: (2024)
por: Shomer, Harry, et al.
Publicado: (2024)
Subgraph Generation for Generalizing on Out-of-Distribution Links
por: Revolinsky, Jay, et al.
Publicado: (2025)
por: Revolinsky, Jay, et al.
Publicado: (2025)
Are Expressive Encoders Necessary for Discrete Graph Generation?
por: Revolinsky, Jay, et al.
Publicado: (2026)
por: Revolinsky, Jay, et al.
Publicado: (2026)
Iterative LLM-Based Generation and Refinement of Distracting Conditions in Math Word Problems
por: Yang, Kaiqi, et al.
Publicado: (2025)
por: Yang, Kaiqi, et al.
Publicado: (2025)
Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem Solutions
por: Li, Hang, et al.
Publicado: (2024)
por: Li, Hang, et al.
Publicado: (2024)
Star-Agents: Automatic Data Optimization with LLM Agents for Instruction Tuning
por: Zhou, Hang, et al.
Publicado: (2024)
por: Zhou, Hang, et al.
Publicado: (2024)
Sub-graph Based Diffusion Model for Link Prediction
por: Li, Hang, et al.
Publicado: (2024)
por: Li, Hang, et al.
Publicado: (2024)
Knowledge Tagging System on Math Questions via LLMs with Flexible Demonstration Retriever
por: Li, Hang, et al.
Publicado: (2024)
por: Li, Hang, et al.
Publicado: (2024)
Empowering GraphRAG with Knowledge Filtering and Integration
por: Guo, Kai, et al.
Publicado: (2025)
por: Guo, Kai, et al.
Publicado: (2025)
Towards Human-Like Grading: A Unified LLM-Enhanced Framework for Subjective Question Evaluation
por: Zhua, Fanwei, et al.
Publicado: (2025)
por: Zhua, Fanwei, et al.
Publicado: (2025)
LLM-Powered Automatic Translation and Urgency in Crisis Scenarios
por: Ticona, Belu, et al.
Publicado: (2026)
por: Ticona, Belu, et al.
Publicado: (2026)
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
por: Ryan, Michael J., et al.
Publicado: (2025)
por: Ryan, Michael J., et al.
Publicado: (2025)
Reasoning by Exploration: A Unified Approach to Retrieval and Generation over Graphs
por: Han, Haoyu, et al.
Publicado: (2025)
por: Han, Haoyu, et al.
Publicado: (2025)
Hey AI Can You Grade My Essay?: Automatic Essay Grading
por: Maliha, Maisha, et al.
Publicado: (2024)
por: Maliha, Maisha, et al.
Publicado: (2024)
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
por: Li, Weiyue, et al.
Publicado: (2026)
por: Li, Weiyue, et al.
Publicado: (2026)
Beyond Static Retrieval: Opportunities and Pitfalls of Iterative Retrieval in GraphRAG
por: Guo, Kai, et al.
Publicado: (2025)
por: Guo, Kai, et al.
Publicado: (2025)
Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models
por: Murthy, Rithesh, et al.
Publicado: (2025)
por: Murthy, Rithesh, et al.
Publicado: (2025)
TALEC: Teach Your LLM to Evaluate in Specific Domain with In-house Criteria by Criteria Division and Zero-shot Plus Few-shot
por: Zhang, Kaiqi, et al.
Publicado: (2024)
por: Zhang, Kaiqi, et al.
Publicado: (2024)
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
por: Wang, Yidong, et al.
Publicado: (2023)
por: Wang, Yidong, et al.
Publicado: (2023)
Enhancing ID and Text Fusion via Alternative Training in Session-based Recommendation
por: Li, Juanhui, et al.
Publicado: (2024)
por: Li, Juanhui, et al.
Publicado: (2024)
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
por: Shao, Yijia, et al.
Publicado: (2024)
por: Shao, Yijia, et al.
Publicado: (2024)
Re-evaluating Automatic LLM System Ranking for Alignment with Human Preference
por: Gao, Mingqi, et al.
Publicado: (2024)
por: Gao, Mingqi, et al.
Publicado: (2024)
Beyond Partisan Leaning: A Comparative Analysis of Political Bias in Large Language Models
por: Peng, Tai-Quan, et al.
Publicado: (2024)
por: Peng, Tai-Quan, et al.
Publicado: (2024)
LegalChainReasoner: A Legal Chain-guided Framework for Criminal Judicial Opinion Generation
por: Shi, Weizhe, et al.
Publicado: (2025)
por: Shi, Weizhe, et al.
Publicado: (2025)
Higher-order Structure Boosts Link Prediction on Temporal Graphs
por: Liu, Jingzhe, et al.
Publicado: (2025)
por: Liu, Jingzhe, et al.
Publicado: (2025)
Empowering Molecule Discovery for Molecule-Caption Translation with Large Language Models: A ChatGPT Perspective
por: Li, Jiatong, et al.
Publicado: (2023)
por: Li, Jiatong, et al.
Publicado: (2023)
GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
por: Li, Xueyi, et al.
Publicado: (2026)
por: Li, Xueyi, et al.
Publicado: (2026)
AutoPDL: Automatic Prompt Optimization for LLM Agents
por: Spiess, Claudio, et al.
Publicado: (2025)
por: Spiess, Claudio, et al.
Publicado: (2025)
Zodiac: A Cardiologist-Level LLM Framework for Multi-Agent Diagnostics
por: Zhou, Yuan, et al.
Publicado: (2024)
por: Zhou, Yuan, et al.
Publicado: (2024)
Automatic Transmission for LLM Tiers: Optimizing Cost and Accuracy in Large Language Models
por: Na, Injae, et al.
Publicado: (2025)
por: Na, Injae, et al.
Publicado: (2025)
Ejemplares similares
-
LLM-based Automated Grading with Human-in-the-Loop
por: Chu, Yucheng, et al.
Publicado: (2025) -
Optimizing In-Context Demonstrations for LLM-based Automated Grading
por: Chu, Yucheng, et al.
Publicado: (2026) -
Confusion-Aware Rubric Optimization for LLM-based Automated Grading
por: Chu, Yucheng, et al.
Publicado: (2026) -
From Flat to Structural: Enhancing Automated Short Answer Grading with GraphRAG
por: Chu, Yucheng, et al.
Publicado: (2026) -
How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment
por: Li, Hang, et al.
Publicado: (2026)