Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, Steve, Junior, Gilberto Titericz, Balough, Tom, Zhou, Wenfei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating Metrics for Safety with LLM-as-Judges
di: Clegg, Kester, et al.
Pubblicazione: (2025)
di: Clegg, Kester, et al.
Pubblicazione: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
di: Shi, Lin, et al.
Pubblicazione: (2024)
di: Shi, Lin, et al.
Pubblicazione: (2024)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
di: Wang, Yidong, et al.
Pubblicazione: (2025)
di: Wang, Yidong, et al.
Pubblicazione: (2025)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
di: Thakur, Aman Singh, et al.
Pubblicazione: (2024)
di: Thakur, Aman Singh, et al.
Pubblicazione: (2024)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
di: Lu, Junyu, et al.
Pubblicazione: (2025)
di: Lu, Junyu, et al.
Pubblicazione: (2025)
A Survey on LLM-as-a-Judge
di: Gu, Jiawei, et al.
Pubblicazione: (2024)
di: Gu, Jiawei, et al.
Pubblicazione: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024)
di: Tan, Sijun, et al.
Pubblicazione: (2024)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
di: Tong, Terry, et al.
Pubblicazione: (2025)
di: Tong, Terry, et al.
Pubblicazione: (2025)
Automated Concept Discovery for LLM-as-a-Judge Preference Analysis
di: Wedgwood, James, et al.
Pubblicazione: (2026)
di: Wedgwood, James, et al.
Pubblicazione: (2026)
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
di: Findeis, Arduin, et al.
Pubblicazione: (2025)
di: Findeis, Arduin, et al.
Pubblicazione: (2025)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
di: Chen, Dongping, et al.
Pubblicazione: (2024)
di: Chen, Dongping, et al.
Pubblicazione: (2024)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
di: Zhou, Xin, et al.
Pubblicazione: (2025)
di: Zhou, Xin, et al.
Pubblicazione: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
di: Jiang, Hongchao, et al.
Pubblicazione: (2025)
di: Jiang, Hongchao, et al.
Pubblicazione: (2025)
Same Verdict, Different Reasons: LLM-as-a-Judge and Clinician Disagreement on Medical Chatbot Completeness
di: DeLucia, Alexandra, et al.
Pubblicazione: (2026)
di: DeLucia, Alexandra, et al.
Pubblicazione: (2026)
LLM-as-a-Judge for Time Series Explanations
di: Sivalingam, Preetham, et al.
Pubblicazione: (2026)
di: Sivalingam, Preetham, et al.
Pubblicazione: (2026)
Counterargument for Critical Thinking as Judged by AI and Humans
di: Adewumi, Tosin, et al.
Pubblicazione: (2026)
di: Adewumi, Tosin, et al.
Pubblicazione: (2026)
EssayJudge: A Multi-Granular Benchmark for Assessing Automated Essay Scoring Capabilities of Multimodal Large Language Models
di: Su, Jiamin, et al.
Pubblicazione: (2025)
di: Su, Jiamin, et al.
Pubblicazione: (2025)
Think-J: Learning to Think for Generative LLM-as-a-Judge
di: Huang, Hui, et al.
Pubblicazione: (2025)
di: Huang, Hui, et al.
Pubblicazione: (2025)
JudgeLRM: Large Reasoning Models as a Judge
di: Chen, Nuo, et al.
Pubblicazione: (2025)
di: Chen, Nuo, et al.
Pubblicazione: (2025)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
di: Barnes, Jeremy, et al.
Pubblicazione: (2025)
di: Barnes, Jeremy, et al.
Pubblicazione: (2025)
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
di: Sun, Xin, et al.
Pubblicazione: (2026)
di: Sun, Xin, et al.
Pubblicazione: (2026)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
di: Yuan, Tongxin, et al.
Pubblicazione: (2024)
di: Yuan, Tongxin, et al.
Pubblicazione: (2024)
How Trustworthy Are LLM-as-Judge Ratings for Interpretive Responses? Implications for Qualitative Research Workflows
di: Han, Songhee, et al.
Pubblicazione: (2026)
di: Han, Songhee, et al.
Pubblicazione: (2026)
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation
di: Muhamed, Aashiq
Pubblicazione: (2025)
di: Muhamed, Aashiq
Pubblicazione: (2025)
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
di: Zhou, Xiaolin, et al.
Pubblicazione: (2026)
di: Zhou, Xiaolin, et al.
Pubblicazione: (2026)
M-Prometheus: A Suite of Open Multilingual LLM Judges
di: Pombal, José, et al.
Pubblicazione: (2025)
di: Pombal, José, et al.
Pubblicazione: (2025)
Are We on the Right Way to Assessing LLM-as-a-Judge?
di: Feng, Yuanning, et al.
Pubblicazione: (2025)
di: Feng, Yuanning, et al.
Pubblicazione: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
di: Bologna, Federica, et al.
Pubblicazione: (2026)
di: Bologna, Federica, et al.
Pubblicazione: (2026)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
di: Ye, Jiayi, et al.
Pubblicazione: (2024)
di: Ye, Jiayi, et al.
Pubblicazione: (2024)
Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters
di: Zhang, Xingjian, et al.
Pubblicazione: (2025)
di: Zhang, Xingjian, et al.
Pubblicazione: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
di: Yang, Langqi, et al.
Pubblicazione: (2025)
di: Yang, Langqi, et al.
Pubblicazione: (2025)
JudgeLM: Fine-tuned Large Language Models are Scalable Judges
di: Zhu, Lianghui, et al.
Pubblicazione: (2023)
di: Zhu, Lianghui, et al.
Pubblicazione: (2023)
EasyJudge: an Easy-to-use Tool for Comprehensive Response Evaluation of LLMs
di: Li, Yijie, et al.
Pubblicazione: (2024)
di: Li, Yijie, et al.
Pubblicazione: (2024)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
di: Yang, Jinming, et al.
Pubblicazione: (2026)
di: Yang, Jinming, et al.
Pubblicazione: (2026)
Verdict: A Library for Scaling Judge-Time Compute
di: Kalra, Nimit, et al.
Pubblicazione: (2025)
di: Kalra, Nimit, et al.
Pubblicazione: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
Agent-as-a-Judge
di: You, Runyang, et al.
Pubblicazione: (2026)
di: You, Runyang, et al.
Pubblicazione: (2026)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
di: Saha, Swarnadeep, et al.
Pubblicazione: (2025)
di: Saha, Swarnadeep, et al.
Pubblicazione: (2025)
Criterion Validity of LLM-as-Judge for Business Outcomes in Conversational Commerce
di: Chen, Liang, et al.
Pubblicazione: (2026)
di: Chen, Liang, et al.
Pubblicazione: (2026)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
di: Yoon, Kanghoon, et al.
Pubblicazione: (2025)
di: Yoon, Kanghoon, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Evaluating Metrics for Safety with LLM-as-Judges
di: Clegg, Kester, et al.
Pubblicazione: (2025) -
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
di: Shi, Lin, et al.
Pubblicazione: (2024) -
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
di: Wang, Yidong, et al.
Pubblicazione: (2025) -
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
di: Thakur, Aman Singh, et al.
Pubblicazione: (2024) -
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
di: Lu, Junyu, et al.
Pubblicazione: (2025)