Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Weiyue, Zhao, Minda, Dong, Weixuan, Cai, Jiahui, Wei, Yuze, Pocress, Michael, Li, Yi, Yuan, Wanyan, Wang, Xiaoyue, Hou, Ruoyu, Lou, Kaiyuan, Zeng, Wenqi, Yang, Yutong, Du, Yilun, Wang, Mengyu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
por: Yoon, Sung-Hoon, et al.
Publicado: (2026)
por: Yoon, Sung-Hoon, et al.
Publicado: (2026)
Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
por: Zhao, Minda, et al.
Publicado: (2026)
por: Zhao, Minda, et al.
Publicado: (2026)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
por: Zhou, Yilun, et al.
Publicado: (2025)
por: Zhou, Yilun, et al.
Publicado: (2025)
TrustTrade: Human-Inspired Selective Consensus Reduces Decision Uncertainty in LLM Trading Agents
por: Li, Minghan, et al.
Publicado: (2026)
por: Li, Minghan, et al.
Publicado: (2026)
DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training
por: Tian, Xiaoyu, et al.
Publicado: (2025)
por: Tian, Xiaoyu, et al.
Publicado: (2025)
Rubric-Conditioned LLM Grading: Alignment, Uncertainty, and Robustness
por: Deng, Haotian, et al.
Publicado: (2025)
por: Deng, Haotian, et al.
Publicado: (2025)
Individual and Combined Effects of English as a Second Language and Typos on LLM Performance
por: Liu, Serena, et al.
Publicado: (2026)
por: Liu, Serena, et al.
Publicado: (2026)
LLM-based Automated Grading with Human-in-the-Loop
por: Chu, Yucheng, et al.
Publicado: (2025)
por: Chu, Yucheng, et al.
Publicado: (2025)
GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
por: Li, Xueyi, et al.
Publicado: (2026)
por: Li, Xueyi, et al.
Publicado: (2026)
Data Caching for Enterprise-Grade Petabyte-Scale OLAP
por: Tang, Chunxu, et al.
Publicado: (2024)
por: Tang, Chunxu, et al.
Publicado: (2024)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
por: Cai, Yunna, et al.
Publicado: (2025)
por: Cai, Yunna, et al.
Publicado: (2025)
Lightweight Multimodal LLM-Enabled Cost-Effective Defect Grading of Power Transmission Equipment
por: Wang, Tao, et al.
Publicado: (2026)
por: Wang, Tao, et al.
Publicado: (2026)
Optimizing In-Context Demonstrations for LLM-based Automated Grading
por: Chu, Yucheng, et al.
Publicado: (2026)
por: Chu, Yucheng, et al.
Publicado: (2026)
LLM Review: Enhancing Creative Writing via Blind Peer Review Feedback
por: Li, Weiyue, et al.
Publicado: (2026)
por: Li, Weiyue, et al.
Publicado: (2026)
Enhancing LLM-Based Short Answer Grading with Retrieval-Augmented Generation
por: Chu, Yucheng, et al.
Publicado: (2025)
por: Chu, Yucheng, et al.
Publicado: (2025)
Assessing GPT Performance in a Proof-Based University-Level Course Under Blind Grading
por: Ding, Ming, et al.
Publicado: (2025)
por: Ding, Ming, et al.
Publicado: (2025)
MedConclusion: A Benchmark for Biomedical Conclusion Generation from Structured Abstracts
por: Li, Weiyue, et al.
Publicado: (2026)
por: Li, Weiyue, et al.
Publicado: (2026)
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
por: Deshpande, Darshan, et al.
Publicado: (2024)
por: Deshpande, Darshan, et al.
Publicado: (2024)
Confusion-Aware Rubric Optimization for LLM-based Automated Grading
por: Chu, Yucheng, et al.
Publicado: (2026)
por: Chu, Yucheng, et al.
Publicado: (2026)
Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory
por: Cong, Longwei, et al.
Publicado: (2026)
por: Cong, Longwei, et al.
Publicado: (2026)
Human-in-the-Loop LLM Grading for Handwritten Mathematics Assessments
por: Vanhoyweghen, Arne, et al.
Publicado: (2026)
por: Vanhoyweghen, Arne, et al.
Publicado: (2026)
TriAlignXA: An Explainable Trilemma Alignment Framework for Trustworthy Agri-product Grading
por: Xie, Jianfei, et al.
Publicado: (2025)
por: Xie, Jianfei, et al.
Publicado: (2025)
Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
por: Ho, Xanh, et al.
Publicado: (2025)
por: Ho, Xanh, et al.
Publicado: (2025)
Towards Consumer-Grade Cybersickness Prediction: Multi-Model Alignment for Real-Time Vision-Only Inference
por: Zhu, Yitong, et al.
Publicado: (2025)
por: Zhu, Yitong, et al.
Publicado: (2025)
A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization
por: Chu, Yucheng, et al.
Publicado: (2024)
por: Chu, Yucheng, et al.
Publicado: (2024)
SketchJudge: A Diagnostic Benchmark for Grading Hand-drawn Diagrams with Multimodal Large Language Models
por: Su, Yuhang, et al.
Publicado: (2026)
por: Su, Yuhang, et al.
Publicado: (2026)
RocketEval: Efficient Automated LLM Evaluation via Grading Checklist
por: Wei, Tianjun, et al.
Publicado: (2025)
por: Wei, Tianjun, et al.
Publicado: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
por: Liu, Yixin, et al.
Publicado: (2025)
por: Liu, Yixin, et al.
Publicado: (2025)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
por: Cheng, Ruoxi, et al.
Publicado: (2025)
por: Cheng, Ruoxi, et al.
Publicado: (2025)
LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment
por: Li, Lingyao, et al.
Publicado: (2026)
por: Li, Lingyao, et al.
Publicado: (2026)
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
por: Chan, Chi-Min, et al.
Publicado: (2025)
por: Chan, Chi-Min, et al.
Publicado: (2025)
Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems
por: Chhabra, Mukul, et al.
Publicado: (2026)
por: Chhabra, Mukul, et al.
Publicado: (2026)
Scaling LLM Inference with Optimized Sample Compute Allocation
por: Zhang, Kexun, et al.
Publicado: (2024)
por: Zhang, Kexun, et al.
Publicado: (2024)
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
por: Wang, Yutong, et al.
Publicado: (2025)
por: Wang, Yutong, et al.
Publicado: (2025)
Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations
por: Lai, Peng, et al.
Publicado: (2025)
por: Lai, Peng, et al.
Publicado: (2025)
MultEval: Supporting Collaborative Alignment for LLM-as-a-Judge Evaluation Criteria
por: Chiang, Charles, et al.
Publicado: (2026)
por: Chiang, Charles, et al.
Publicado: (2026)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
por: Yang, Bo, et al.
Publicado: (2026)
por: Yang, Bo, et al.
Publicado: (2026)
Core-Elements for Large-Scale Least Squares Estimation
por: Li, Mengyu, et al.
Publicado: (2022)
por: Li, Mengyu, et al.
Publicado: (2022)
LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data
por: Meisenbacher, Stephen, et al.
Publicado: (2025)
por: Meisenbacher, Stephen, et al.
Publicado: (2025)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
por: Wei, Hui, et al.
Publicado: (2024)
por: Wei, Hui, et al.
Publicado: (2024)
Ejemplares similares
-
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
por: Yoon, Sung-Hoon, et al.
Publicado: (2026) -
Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
por: Zhao, Minda, et al.
Publicado: (2026) -
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
por: Zhou, Yilun, et al.
Publicado: (2025) -
TrustTrade: Human-Inspired Selective Consensus Reduces Decision Uncertainty in LLM Trading Agents
por: Li, Minghan, et al.
Publicado: (2026) -
DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training
por: Tian, Xiaoyu, et al.
Publicado: (2025)