Salvato in:
| Autori principali: | Zhu, Ziyi, Tieleman, Olivier, Bukhtiyarov, Alexey, Chen, Jinghong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2603.01865 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating Scoring Bias in LLM-as-a-Judge
di: Li, Qingquan, et al.
Pubblicazione: (2025)
di: Li, Qingquan, et al.
Pubblicazione: (2025)
Quantifying and Mitigating Self-Preference Bias of LLM Judges
di: Yang, Jinming, et al.
Pubblicazione: (2026)
di: Yang, Jinming, et al.
Pubblicazione: (2026)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
di: Marioriyad, Arash, et al.
Pubblicazione: (2025)
di: Marioriyad, Arash, et al.
Pubblicazione: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
di: Shi, Lin, et al.
Pubblicazione: (2024)
di: Shi, Lin, et al.
Pubblicazione: (2024)
Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judge
di: Liu, Zhuo, et al.
Pubblicazione: (2025)
di: Liu, Zhuo, et al.
Pubblicazione: (2025)
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization
di: Zhou, Hongli, et al.
Pubblicazione: (2026)
di: Zhou, Hongli, et al.
Pubblicazione: (2026)
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
di: Li, Haitao, et al.
Pubblicazione: (2024)
di: Li, Haitao, et al.
Pubblicazione: (2024)
Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge
di: Fujinuma, Yoshinari
Pubblicazione: (2025)
di: Fujinuma, Yoshinari
Pubblicazione: (2025)
DIAL: Direct Iterative Adversarial Learning for Realistic Multi-Turn Dialogue Simulation
di: Zhu, Ziyi, et al.
Pubblicazione: (2025)
di: Zhu, Ziyi, et al.
Pubblicazione: (2025)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
di: Lai, Peng, et al.
Pubblicazione: (2026)
di: Lai, Peng, et al.
Pubblicazione: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024)
di: Tan, Sijun, et al.
Pubblicazione: (2024)
Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information Bottleneck
di: Zhang, Hongbin, et al.
Pubblicazione: (2026)
di: Zhang, Hongbin, et al.
Pubblicazione: (2026)
Self-Preference Bias in LLM-as-a-Judge
di: Wataoka, Koki, et al.
Pubblicazione: (2024)
di: Wataoka, Koki, et al.
Pubblicazione: (2024)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
di: Belmadani, Ikram, et al.
Pubblicazione: (2026)
di: Belmadani, Ikram, et al.
Pubblicazione: (2026)
CalibraEval: Calibrating Prediction Distribution to Mitigate Selection Bias in LLMs-as-Judges
di: Li, Haitao, et al.
Pubblicazione: (2024)
di: Li, Haitao, et al.
Pubblicazione: (2024)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
di: Thakur, Aman Singh, et al.
Pubblicazione: (2024)
di: Thakur, Aman Singh, et al.
Pubblicazione: (2024)
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
di: Lee, Sua, et al.
Pubblicazione: (2026)
di: Lee, Sua, et al.
Pubblicazione: (2026)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
di: Yang, Bo, et al.
Pubblicazione: (2026)
di: Yang, Bo, et al.
Pubblicazione: (2026)
Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges
di: Tang, Yuqi, et al.
Pubblicazione: (2025)
di: Tang, Yuqi, et al.
Pubblicazione: (2025)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
di: Lee, Dongryeol, et al.
Pubblicazione: (2026)
di: Lee, Dongryeol, et al.
Pubblicazione: (2026)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
di: Clegg, Kester, et al.
Pubblicazione: (2025)
di: Clegg, Kester, et al.
Pubblicazione: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
di: Tong, Terry, et al.
Pubblicazione: (2025)
di: Tong, Terry, et al.
Pubblicazione: (2025)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
di: Wang, Yidong, et al.
Pubblicazione: (2025)
di: Wang, Yidong, et al.
Pubblicazione: (2025)
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
di: Zhou, Xiaolin, et al.
Pubblicazione: (2026)
di: Zhou, Xiaolin, et al.
Pubblicazione: (2026)
Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges
di: Tapwal, Riya, et al.
Pubblicazione: (2026)
di: Tapwal, Riya, et al.
Pubblicazione: (2026)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
di: Moon, Jiwon, et al.
Pubblicazione: (2025)
di: Moon, Jiwon, et al.
Pubblicazione: (2025)
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
di: Marioriyad, Arash, et al.
Pubblicazione: (2026)
di: Marioriyad, Arash, et al.
Pubblicazione: (2026)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
di: Chen, Junjie, et al.
Pubblicazione: (2026)
di: Chen, Junjie, et al.
Pubblicazione: (2026)
Quantitative LLM Judges
di: Sahoo, Aishwarya, et al.
Pubblicazione: (2025)
di: Sahoo, Aishwarya, et al.
Pubblicazione: (2025)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
di: Huang, Hui, et al.
Pubblicazione: (2024)
di: Huang, Hui, et al.
Pubblicazione: (2024)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
di: Jiang, Hongchao, et al.
Pubblicazione: (2025)
di: Jiang, Hongchao, et al.
Pubblicazione: (2025)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
di: Bellibatlu, Rohith Reddy, et al.
Pubblicazione: (2026)
di: Bellibatlu, Rohith Reddy, et al.
Pubblicazione: (2026)
Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
di: Liu, Shuliang, et al.
Pubblicazione: (2025)
di: Liu, Shuliang, et al.
Pubblicazione: (2025)
MR. Judge: Multimodal Reasoner as a Judge
di: Pi, Renjie, et al.
Pubblicazione: (2025)
di: Pi, Renjie, et al.
Pubblicazione: (2025)
Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based Evaluation
di: Lee, Dongryeol, et al.
Pubblicazione: (2024)
di: Lee, Dongryeol, et al.
Pubblicazione: (2024)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
di: Wang, Qian, et al.
Pubblicazione: (2025)
di: Wang, Qian, et al.
Pubblicazione: (2025)
Cross-Lingual LLM-Judge Transfer via Evaluation Decomposition
di: Sheth, Ivaxi, et al.
Pubblicazione: (2026)
di: Sheth, Ivaxi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Evaluating Scoring Bias in LLM-as-a-Judge
di: Li, Qingquan, et al.
Pubblicazione: (2025) -
Quantifying and Mitigating Self-Preference Bias of LLM Judges
di: Yang, Jinming, et al.
Pubblicazione: (2026) -
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
di: Marioriyad, Arash, et al.
Pubblicazione: (2025) -
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
di: Shi, Lin, et al.
Pubblicazione: (2024) -
Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judge
di: Liu, Zhuo, et al.
Pubblicazione: (2025)