Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Schroeder, Kayla, Wood-Doughty, Zach |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Reliability of Topic Modeling
par: Schroeder, Kayla, et autres
Publié: (2024)
par: Schroeder, Kayla, et autres
Publié: (2024)
Improving LLM-as-a-Judge Inference with the Judgment Distribution
par: Wang, Victor, et autres
Publié: (2025)
par: Wang, Victor, et autres
Publié: (2025)
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
par: Hwang, Yerin, et autres
Publié: (2025)
par: Hwang, Yerin, et autres
Publié: (2025)
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
par: Wang, Leyao, et autres
Publié: (2026)
par: Wang, Leyao, et autres
Publié: (2026)
How Reliable is Multilingual LLM-as-a-Judge?
par: Fu, Xiyan, et autres
Publié: (2025)
par: Fu, Xiyan, et autres
Publié: (2025)
Can LLM be a Personalized Judge?
par: Dong, Yijiang River, et autres
Publié: (2024)
par: Dong, Yijiang River, et autres
Publié: (2024)
Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge
par: Chen, Luyu, et autres
Publié: (2025)
par: Chen, Luyu, et autres
Publié: (2025)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
par: Wang, Yidong, et autres
Publié: (2025)
par: Wang, Yidong, et autres
Publié: (2025)
Can We Trust LLM Detectors?
par: Sandhan, Jivnesh, et autres
Publié: (2026)
par: Sandhan, Jivnesh, et autres
Publié: (2026)
VERT: Reliable LLM Judges for Radiology Report Evaluation
par: Bologna, Federica, et autres
Publié: (2026)
par: Bologna, Federica, et autres
Publié: (2026)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
par: Jung, Jaehun, et autres
Publié: (2024)
par: Jung, Jaehun, et autres
Publié: (2024)
LLM-REVal: Can We Trust LLM Reviewers Yet?
par: Li, Rui, et autres
Publié: (2025)
par: Li, Rui, et autres
Publié: (2025)
AI vs. Human Judgment of Content Moderation: LLM-as-a-Judge and Ethics-Based Response Refusals
par: Pasch, Stefan
Publié: (2025)
par: Pasch, Stefan
Publié: (2025)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
par: Son, Guijin, et autres
Publié: (2024)
par: Son, Guijin, et autres
Publié: (2024)
SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?
par: Chen, Jiamin, et autres
Publié: (2026)
par: Chen, Jiamin, et autres
Publié: (2026)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
par: Badawi, Abeer, et autres
Publié: (2025)
par: Badawi, Abeer, et autres
Publié: (2025)
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
par: Yamauchi, Yusuke, et autres
Publié: (2025)
par: Yamauchi, Yusuke, et autres
Publié: (2025)
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
par: Sun, Xin, et autres
Publié: (2026)
par: Sun, Xin, et autres
Publié: (2026)
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
par: Lin, Wei-Hsiang, et autres
Publié: (2025)
par: Lin, Wei-Hsiang, et autres
Publié: (2025)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
par: Marioriyad, Arash, et autres
Publié: (2025)
par: Marioriyad, Arash, et autres
Publié: (2025)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
par: Yang, Bo, et autres
Publié: (2026)
par: Yang, Bo, et autres
Publié: (2026)
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
par: Findeis, Arduin, et autres
Publié: (2025)
par: Findeis, Arduin, et autres
Publié: (2025)
Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters
par: Zhang, Xingjian, et autres
Publié: (2025)
par: Zhang, Xingjian, et autres
Publié: (2025)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
par: Gupta, Manan, et autres
Publié: (2026)
par: Gupta, Manan, et autres
Publié: (2026)
Beyond the Surface: Measuring Self-Preference in LLM Judgments
par: Chen, Zhi-Yuan, et autres
Publié: (2025)
par: Chen, Zhi-Yuan, et autres
Publié: (2025)
When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment
par: Ferrer, Robinson, et autres
Publié: (2026)
par: Ferrer, Robinson, et autres
Publié: (2026)
Quantitative LLM Judges
par: Sahoo, Aishwarya, et autres
Publié: (2025)
par: Sahoo, Aishwarya, et autres
Publié: (2025)
Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal Style
par: Baumler, Connor, et autres
Publié: (2026)
par: Baumler, Connor, et autres
Publié: (2026)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
par: Belmadani, Ikram, et autres
Publié: (2026)
par: Belmadani, Ikram, et autres
Publié: (2026)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
par: Schwinn, Leo, et autres
Publié: (2026)
par: Schwinn, Leo, et autres
Publié: (2026)
EnsemJudge: Enhancing Reliability in Chinese LLM-Generated Text Detection through Diverse Model Ensembles
par: Wang, Zhuoshang, et autres
Publié: (2026)
par: Wang, Zhuoshang, et autres
Publié: (2026)
Self-Preference Bias in LLM-as-a-Judge
par: Wataoka, Koki, et autres
Publié: (2024)
par: Wataoka, Koki, et autres
Publié: (2024)
The Necessity of Setting Temperature in LLM-as-a-Judge
par: Li, Lujun, et autres
Publié: (2026)
par: Li, Lujun, et autres
Publié: (2026)
Evaluating Scoring Bias in LLM-as-a-Judge
par: Li, Qingquan, et autres
Publié: (2025)
par: Li, Qingquan, et autres
Publié: (2025)
Benchmarking LLM-based Relevance Judgment Methods
par: Arabzadeh, Negar, et autres
Publié: (2025)
par: Arabzadeh, Negar, et autres
Publié: (2025)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
par: Huang, Hui, et autres
Publié: (2024)
par: Huang, Hui, et autres
Publié: (2024)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
par: Shi, Lin, et autres
Publié: (2024)
par: Shi, Lin, et autres
Publié: (2024)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
par: Bellibatlu, Rohith Reddy, et autres
Publié: (2026)
par: Bellibatlu, Rohith Reddy, et autres
Publié: (2026)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
par: Hong, Yihan, et autres
Publié: (2026)
par: Hong, Yihan, et autres
Publié: (2026)
A Survey on LLM-as-a-Judge
par: Gu, Jiawei, et autres
Publié: (2024)
par: Gu, Jiawei, et autres
Publié: (2024)
Documents similaires
-
Reliability of Topic Modeling
par: Schroeder, Kayla, et autres
Publié: (2024) -
Improving LLM-as-a-Judge Inference with the Judgment Distribution
par: Wang, Victor, et autres
Publié: (2025) -
Can You Trick the Grader? Adversarial Persuasion of LLM Judges
par: Hwang, Yerin, et autres
Publié: (2025) -
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
par: Wang, Leyao, et autres
Publié: (2026) -
How Reliable is Multilingual LLM-as-a-Judge?
par: Fu, Xiyan, et autres
Publié: (2025)