Evaluating Metrics for Safety with LLM-as-Judges
Fuente:
arXiv
Saved in:
| Main Authors: | Clegg, Kester, Hawkins, Richard, Habli, Ibrahim, Lawton, Tom |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
by: Sun, Bian, et al.
Published: (2026)
by: Sun, Bian, et al.
Published: (2026)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
by: Han, Steve, et al.
Published: (2025)
by: Han, Steve, et al.
Published: (2025)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
by: Alam, Firoj, et al.
Published: (2026)
by: Alam, Firoj, et al.
Published: (2026)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
by: Yuan, Tongxin, et al.
Published: (2024)
by: Yuan, Tongxin, et al.
Published: (2024)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
by: Barnes, Jeremy, et al.
Published: (2025)
by: Barnes, Jeremy, et al.
Published: (2025)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
by: Collot, Stephane, et al.
Published: (2025)
by: Collot, Stephane, et al.
Published: (2025)
Upstream and Downstream AI Safety: Both on the Same River?
by: McDermid, John, et al.
Published: (2024)
by: McDermid, John, et al.
Published: (2024)
VERT: Reliable LLM Judges for Radiology Report Evaluation
by: Bologna, Federica, et al.
Published: (2026)
by: Bologna, Federica, et al.
Published: (2026)
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
by: Reese, May Lynn, et al.
Published: (2026)
by: Reese, May Lynn, et al.
Published: (2026)
Clear, Compelling Arguments: Rethinking the Foundations of Frontier AI Safety Cases
by: Feakins, Shaun, et al.
Published: (2026)
by: Feakins, Shaun, et al.
Published: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
by: Schwinn, Leo, et al.
Published: (2026)
by: Schwinn, Leo, et al.
Published: (2026)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
by: Thakur, Aman Singh, et al.
Published: (2024)
by: Thakur, Aman Singh, et al.
Published: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
by: Saha, Swarnadeep, et al.
Published: (2025)
by: Saha, Swarnadeep, et al.
Published: (2025)
Learning Run-time Safety Monitors for Machine Learning Components
by: Vardal, Ozan, et al.
Published: (2024)
by: Vardal, Ozan, et al.
Published: (2024)
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
by: Findeis, Arduin, et al.
Published: (2025)
by: Findeis, Arduin, et al.
Published: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
by: Shi, Lin, et al.
Published: (2024)
by: Shi, Lin, et al.
Published: (2024)
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
by: Enguehard, Joseph, et al.
Published: (2025)
by: Enguehard, Joseph, et al.
Published: (2025)
Improving AGI Evaluation: A Data Science Perspective
by: Hawkins, John
Published: (2025)
by: Hawkins, John
Published: (2025)
Exploring Human-AI Complementarity in CPS Diagnosis Using Unimodal and Multimodal BERT Models
by: Wong, Kester, et al.
Published: (2025)
by: Wong, Kester, et al.
Published: (2025)
Explainable Collaborative Problem Solving Diagnosis with BERT using SHAP and its Implications for Teacher Adoption
by: Wong, Kester, et al.
Published: (2025)
by: Wong, Kester, et al.
Published: (2025)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
by: Shi, Zhichao, et al.
Published: (2025)
by: Shi, Zhichao, et al.
Published: (2025)
Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
by: Verga, Pat, et al.
Published: (2024)
by: Verga, Pat, et al.
Published: (2024)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
by: Jin, Jiho, et al.
Published: (2026)
by: Jin, Jiho, et al.
Published: (2026)
Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems
by: Chhabra, Mukul, et al.
Published: (2026)
by: Chhabra, Mukul, et al.
Published: (2026)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
by: Wang, Yidong, et al.
Published: (2025)
by: Wang, Yidong, et al.
Published: (2025)
A Survey on LLM-as-a-Judge
by: Gu, Jiawei, et al.
Published: (2024)
by: Gu, Jiawei, et al.
Published: (2024)
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
by: Lin, Wei-Hsiang, et al.
Published: (2025)
by: Lin, Wei-Hsiang, et al.
Published: (2025)
LLM-RadJudge: Achieving Radiologist-Level Evaluation for X-Ray Report Generation
by: Wang, Zilong, et al.
Published: (2024)
by: Wang, Zilong, et al.
Published: (2024)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
by: Xu, Austin, et al.
Published: (2025)
by: Xu, Austin, et al.
Published: (2025)
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools
by: Park, Jung In, et al.
Published: (2024)
by: Park, Jung In, et al.
Published: (2024)
LLM-as-a-Judge for Time Series Explanations
by: Sivalingam, Preetham, et al.
Published: (2026)
by: Sivalingam, Preetham, et al.
Published: (2026)
EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
by: Mohammadi, Hadi, et al.
Published: (2025)
by: Mohammadi, Hadi, et al.
Published: (2025)
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
by: Gisserot-Boukhlef, Hippolyte, et al.
Published: (2026)
by: Gisserot-Boukhlef, Hippolyte, et al.
Published: (2026)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
by: Cho, Yousang, et al.
Published: (2025)
by: Cho, Yousang, et al.
Published: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
by: Tong, Terry, et al.
Published: (2025)
by: Tong, Terry, et al.
Published: (2025)
Are We on the Right Way to Assessing LLM-as-a-Judge?
by: Feng, Yuanning, et al.
Published: (2025)
by: Feng, Yuanning, et al.
Published: (2025)
Similar Items
-
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025) -
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
by: Sun, Bian, et al.
Published: (2026) -
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
by: Zhou, Xin, et al.
Published: (2025) -
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
by: Han, Steve, et al.
Published: (2025) -
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
by: Alam, Firoj, et al.
Published: (2026)