An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Hui, Bu, Xingyuan, Zhou, Hongli, Qu, Yingqi, Liu, Jing, Yang, Muyun, Xu, Bing, Zhao, Tiejun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
von: Zhou, Hongli, et al.
Veröffentlicht: (2026)
von: Zhou, Hongli, et al.
Veröffentlicht: (2026)
RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
von: Zhou, Hongli, et al.
Veröffentlicht: (2026)
von: Zhou, Hongli, et al.
Veröffentlicht: (2026)
Self-Evaluation of Large Language Model based on Glass-box Features
von: Huang, Hui, et al.
Veröffentlicht: (2024)
von: Huang, Hui, et al.
Veröffentlicht: (2024)
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
von: Huang, Hui, et al.
Veröffentlicht: (2026)
von: Huang, Hui, et al.
Veröffentlicht: (2026)
Think-J: Learning to Think for Generative LLM-as-a-Judge
von: Huang, Hui, et al.
Veröffentlicht: (2025)
von: Huang, Hui, et al.
Veröffentlicht: (2025)
Mitigating the Bias of Large Language Model Evaluation
von: Zhou, Hongli, et al.
Veröffentlicht: (2024)
von: Zhou, Hongli, et al.
Veröffentlicht: (2024)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
von: Shi, Lin, et al.
Veröffentlicht: (2024)
von: Shi, Lin, et al.
Veröffentlicht: (2024)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
von: Yamauchi, Yusuke, et al.
Veröffentlicht: (2025)
von: Yamauchi, Yusuke, et al.
Veröffentlicht: (2025)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
TRACT: Regression-Aware Fine-tuning Meets Chain-of-Thought Reasoning for LLM-as-a-Judge
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2025)
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2025)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
von: Belmadani, Ikram, et al.
Veröffentlicht: (2026)
von: Belmadani, Ikram, et al.
Veröffentlicht: (2026)
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines
von: Soumik, Sadman Kabir
Veröffentlicht: (2026)
von: Soumik, Sadman Kabir
Veröffentlicht: (2026)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
JudgeLM: Fine-tuned Large Language Models are Scalable Judges
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
LLM-as-a-Fuzzy-Judge: Fine-Tuning Large Language Models as a Clinical Evaluation Judge with Fuzzy Logic
von: Zheng, Weibing, et al.
Veröffentlicht: (2025)
von: Zheng, Weibing, et al.
Veröffentlicht: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
von: Li, Xiaochuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaochuan, et al.
Veröffentlicht: (2025)
Evaluating Scoring Bias in LLM-as-a-Judge
von: Li, Qingquan, et al.
Veröffentlicht: (2025)
von: Li, Qingquan, et al.
Veröffentlicht: (2025)
Judging the Judges: A Collection of LLM-Generated Relevance Judgements
von: Rahmani, Hossein A., et al.
Veröffentlicht: (2025)
von: Rahmani, Hossein A., et al.
Veröffentlicht: (2025)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
von: Zhu, Ziyi, et al.
Veröffentlicht: (2026)
von: Zhu, Ziyi, et al.
Veröffentlicht: (2026)
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
von: Eiras, Francisco, et al.
Veröffentlicht: (2025)
von: Eiras, Francisco, et al.
Veröffentlicht: (2025)
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
von: Marioriyad, Arash, et al.
Veröffentlicht: (2025)
von: Marioriyad, Arash, et al.
Veröffentlicht: (2025)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
von: Yang, Bo, et al.
Veröffentlicht: (2026)
von: Yang, Bo, et al.
Veröffentlicht: (2026)
LLM-as-Judge on a Budget
von: Saha, Aadirupa, et al.
Veröffentlicht: (2026)
von: Saha, Aadirupa, et al.
Veröffentlicht: (2026)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
von: Chen, Junjie, et al.
Veröffentlicht: (2026)
von: Chen, Junjie, et al.
Veröffentlicht: (2026)
Who's Your Judge? On the Detectability of LLM-Generated Judgments
von: Li, Dawei, et al.
Veröffentlicht: (2025)
von: Li, Dawei, et al.
Veröffentlicht: (2025)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
How to Correctly Report LLM-as-a-Judge Evaluations
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
Efficient Inference for Noisy LLM-as-a-Judge Evaluation
von: Chen, Yiqun T, et al.
Veröffentlicht: (2026)
von: Chen, Yiqun T, et al.
Veröffentlicht: (2026)
Quantitative LLM Judges
von: Sahoo, Aishwarya, et al.
Veröffentlicht: (2025)
von: Sahoo, Aishwarya, et al.
Veröffentlicht: (2025)
On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization
von: Singh, Janvijay, et al.
Veröffentlicht: (2025)
von: Singh, Janvijay, et al.
Veröffentlicht: (2025)
Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction
von: Sheng, Huanxin, et al.
Veröffentlicht: (2025)
von: Sheng, Huanxin, et al.
Veröffentlicht: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
von: Zhou, Hongli, et al.
Veröffentlicht: (2026) -
RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
von: Zhou, Hongli, et al.
Veröffentlicht: (2026) -
Self-Evaluation of Large Language Model based on Glass-box Features
von: Huang, Hui, et al.
Veröffentlicht: (2024) -
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
von: Huang, Hui, et al.
Veröffentlicht: (2026) -
Think-J: Learning to Think for Generative LLM-as-a-Judge
von: Huang, Hui, et al.
Veröffentlicht: (2025)