Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Hui, He, Shenghua, Xia, Tian, Liu, Fei, Wong, Andy, Lin, Jingyang, Han, Mei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoning
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
von: Lin, Jingyang, et al.
Veröffentlicht: (2025)
PlanGenLLMs: A Modern Survey of LLM Planning Capabilities
von: Wei, Hui, et al.
Veröffentlicht: (2025)
von: Wei, Hui, et al.
Veröffentlicht: (2025)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types
von: Mou, Yutao, et al.
Veröffentlicht: (2024)
von: Mou, Yutao, et al.
Veröffentlicht: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Agri-CPJ: A Training-Free Explainable Framework for Agricultural Pest Diagnosis Using Caption-Prompt-Judge and LLM-as-a-Judge
von: Zhang, Wentao, et al.
Veröffentlicht: (2026)
von: Zhang, Wentao, et al.
Veröffentlicht: (2026)
CPJ: Explainable Agricultural Pest Diagnosis via Caption-Prompt-Judge with LLM-Judged Refinement
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
von: Zhang, Wentao, et al.
Veröffentlicht: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
von: Shi, Lin, et al.
Veröffentlicht: (2024)
von: Shi, Lin, et al.
Veröffentlicht: (2024)
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
von: He, Shenghua, et al.
Veröffentlicht: (2025)
von: He, Shenghua, et al.
Veröffentlicht: (2025)
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
von: Eigler, Lukáš, et al.
Veröffentlicht: (2026)
von: Eigler, Lukáš, et al.
Veröffentlicht: (2026)
LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
von: Bellibatlu, Rohith Reddy, et al.
Veröffentlicht: (2026)
von: Bellibatlu, Rohith Reddy, et al.
Veröffentlicht: (2026)
Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge
von: Wu, Junjie, et al.
Veröffentlicht: (2026)
von: Wu, Junjie, et al.
Veröffentlicht: (2026)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
von: Huang, Hui, et al.
Veröffentlicht: (2024)
von: Huang, Hui, et al.
Veröffentlicht: (2024)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
von: Verga, Pat, et al.
Veröffentlicht: (2024)
von: Verga, Pat, et al.
Veröffentlicht: (2024)
PromptOptMe: Error-Aware Prompt Compression for LLM-based MT Evaluation Metrics
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
von: Chang, Hwan, et al.
Veröffentlicht: (2025)
von: Chang, Hwan, et al.
Veröffentlicht: (2025)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
von: Benjamin, Victoria, et al.
Veröffentlicht: (2024)
von: Benjamin, Victoria, et al.
Veröffentlicht: (2024)
Prompt Attack Detection with LLM-as-a-Judge and Mixture-of-Models
von: Le, Hieu Xuan, et al.
Veröffentlicht: (2026)
von: Le, Hieu Xuan, et al.
Veröffentlicht: (2026)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
von: Sun, Bian, et al.
Veröffentlicht: (2026)
von: Sun, Bian, et al.
Veröffentlicht: (2026)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
von: Zhu, Ziyi, et al.
Veröffentlicht: (2026)
von: Zhu, Ziyi, et al.
Veröffentlicht: (2026)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
Evaluating Scoring Bias in LLM-as-a-Judge
von: Li, Qingquan, et al.
Veröffentlicht: (2025)
von: Li, Qingquan, et al.
Veröffentlicht: (2025)
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
von: Lin, Wei-Hsiang, et al.
Veröffentlicht: (2025)
von: Lin, Wei-Hsiang, et al.
Veröffentlicht: (2025)
Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction
von: Sheng, Huanxin, et al.
Veröffentlicht: (2025)
von: Sheng, Huanxin, et al.
Veröffentlicht: (2025)
Think-J: Learning to Think for Generative LLM-as-a-Judge
von: Huang, Hui, et al.
Veröffentlicht: (2025)
von: Huang, Hui, et al.
Veröffentlicht: (2025)
Transferable Modeling Strategies for Low-Resource LLM Tasks: A Prompt and Alignment-Based Approach
von: Lyu, Shuangquan, et al.
Veröffentlicht: (2025)
von: Lyu, Shuangquan, et al.
Veröffentlicht: (2025)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
von: Wu, Tianhao, et al.
Veröffentlicht: (2024)
Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
von: Maloyan, Narek, et al.
Veröffentlicht: (2025)
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation
von: Zhang, Xinran
Veröffentlicht: (2026)
von: Zhang, Xinran
Veröffentlicht: (2026)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
von: Belmadani, Ikram, et al.
Veröffentlicht: (2026)
von: Belmadani, Ikram, et al.
Veröffentlicht: (2026)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
von: Cai, Yunna, et al.
Veröffentlicht: (2025)
von: Cai, Yunna, et al.
Veröffentlicht: (2025)
CommunityBench: Benchmarking Community-Level Alignment across Diverse Groups and Tasks
von: Lin, Jiayu, et al.
Veröffentlicht: (2026)
von: Lin, Jiayu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoning
von: Lin, Jingyang, et al.
Veröffentlicht: (2025) -
PlanGenLLMs: A Modern Survey of LLM Planning Capabilities
von: Wei, Hui, et al.
Veröffentlicht: (2025) -
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
von: Zhou, Xin, et al.
Veröffentlicht: (2025) -
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025) -
SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types
von: Mou, Yutao, et al.
Veröffentlicht: (2024)