Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Soumik, Sadman Kabir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
von: Shi, Lin, et al.
Veröffentlicht: (2024)
von: Shi, Lin, et al.
Veröffentlicht: (2024)
Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations
von: Jain, Suryaansh, et al.
Veröffentlicht: (2025)
von: Jain, Suryaansh, et al.
Veröffentlicht: (2025)
To Judge or not to Judge: Using LLM Judgements for Advertiser Keyphrase Relevance at eBay
von: Dey, Soumik, et al.
Veröffentlicht: (2025)
von: Dey, Soumik, et al.
Veröffentlicht: (2025)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
von: Yang, Jinming, et al.
Veröffentlicht: (2026)
von: Yang, Jinming, et al.
Veröffentlicht: (2026)
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
von: Fujinuma, Yoshinari
Veröffentlicht: (2025)
von: Fujinuma, Yoshinari
Veröffentlicht: (2025)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
Towards Provably Unbiased LLM Judges via Bias-Bounded Evaluation
von: Feuer, Benjamin, et al.
Veröffentlicht: (2026)
von: Feuer, Benjamin, et al.
Veröffentlicht: (2026)
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
von: Lee, Sua, et al.
Veröffentlicht: (2026)
von: Lee, Sua, et al.
Veröffentlicht: (2026)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
von: Li, Xiaochuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaochuan, et al.
Veröffentlicht: (2025)
LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection
von: Hossain, Akram, et al.
Veröffentlicht: (2026)
von: Hossain, Akram, et al.
Veröffentlicht: (2026)
Judge Reliability Harness: Stress Testing the Reliability of LLM Judges
von: Dev, Sunishchal, et al.
Veröffentlicht: (2026)
von: Dev, Sunishchal, et al.
Veröffentlicht: (2026)
Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
von: Lai, Peng, et al.
Veröffentlicht: (2026)
von: Lai, Peng, et al.
Veröffentlicht: (2026)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
von: Wang, Yidong, et al.
Veröffentlicht: (2025)
Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering
von: Zhao, Zixiao, et al.
Veröffentlicht: (2026)
von: Zhao, Zixiao, et al.
Veröffentlicht: (2026)
When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
von: Yu, Fangyi
Veröffentlicht: (2025)
von: Yu, Fangyi
Veröffentlicht: (2025)
Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplifications and Resistance in Multi-Agent Based LLM-as-Judge
von: Ma, Chiyu, et al.
Veröffentlicht: (2025)
von: Ma, Chiyu, et al.
Veröffentlicht: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
A Survey on LLM-as-a-Judge
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
von: Zhou, Xiaolin, et al.
Veröffentlicht: (2026)
von: Zhou, Xiaolin, et al.
Veröffentlicht: (2026)
Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons
von: Hu, Renjun, et al.
Veröffentlicht: (2025)
von: Hu, Renjun, et al.
Veröffentlicht: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
von: Han, Steve, et al.
Veröffentlicht: (2025)
von: Han, Steve, et al.
Veröffentlicht: (2025)
JudgeLRM: Large Reasoning Models as a Judge
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K--12 Science Instructional Materials
von: He, Peng, et al.
Veröffentlicht: (2026)
von: He, Peng, et al.
Veröffentlicht: (2026)
JudgeFlow: Agentic Workflow Optimization via Block Judge
von: Ma, Zihan, et al.
Veröffentlicht: (2026)
von: Ma, Zihan, et al.
Veröffentlicht: (2026)
VERT: Reliable LLM Judges for Radiology Report Evaluation
von: Bologna, Federica, et al.
Veröffentlicht: (2026)
von: Bologna, Federica, et al.
Veröffentlicht: (2026)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2025)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
von: Spiliopoulou, Evangelia, et al.
Veröffentlicht: (2025)
von: Spiliopoulou, Evangelia, et al.
Veröffentlicht: (2025)
Agent-as-a-Judge: Evaluate Agents with Agents
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2024)
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2024)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
von: Zhu, Ziyi, et al.
Veröffentlicht: (2026)
von: Zhu, Ziyi, et al.
Veröffentlicht: (2026)
Auto-Prompt Ensemble for LLM Judge
von: Li, Jiajie, et al.
Veröffentlicht: (2025)
von: Li, Jiajie, et al.
Veröffentlicht: (2025)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
PentestJudge: Judging Agent Behavior Against Operational Requirements
von: Caldwell, Shane, et al.
Veröffentlicht: (2025)
von: Caldwell, Shane, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
von: Shi, Lin, et al.
Veröffentlicht: (2024) -
Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations
von: Jain, Suryaansh, et al.
Veröffentlicht: (2025) -
To Judge or not to Judge: Using LLM Judgements for Advertiser Keyphrase Relevance at eBay
von: Dey, Soumik, et al.
Veröffentlicht: (2025) -
Quantifying and Mitigating Self-Preference Bias of LLM Judges
von: Yang, Jinming, et al.
Veröffentlicht: (2026) -
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
von: Fujinuma, Yoshinari
Veröffentlicht: (2025)