How Sensitive Are Safety Benchmarks to Judge Configuration Choices?
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Zhang, Xinran |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
par: Bellibatlu, Rohith Reddy, et autres
Publié: (2026)
par: Bellibatlu, Rohith Reddy, et autres
Publié: (2026)
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation
par: Zhang, Xinran
Publié: (2026)
par: Zhang, Xinran
Publié: (2026)
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
par: Yamauchi, Yusuke, et autres
Publié: (2025)
par: Yamauchi, Yusuke, et autres
Publié: (2025)
Beyond Creed: A Non-Identity Safety Condition A Strong Empirical Alternative to Identity Framing in Low-Data LoRA Fine-Tuning
par: Zhang, Xinran
Publié: (2026)
par: Zhang, Xinran
Publié: (2026)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
par: Yuan, Tongxin, et autres
Publié: (2024)
par: Yuan, Tongxin, et autres
Publié: (2024)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
par: Wang, Yidong, et autres
Publié: (2025)
par: Wang, Yidong, et autres
Publié: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
par: Clegg, Kester, et autres
Publié: (2025)
par: Clegg, Kester, et autres
Publié: (2025)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
par: Tang, Zhenwei, et autres
Publié: (2026)
par: Tang, Zhenwei, et autres
Publié: (2026)
XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
par: Choi, Dasol, et autres
Publié: (2026)
par: Choi, Dasol, et autres
Publié: (2026)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
par: Chen, Junjie, et autres
Publié: (2026)
par: Chen, Junjie, et autres
Publié: (2026)
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
par: Wen, Bosi, et autres
Publié: (2026)
par: Wen, Bosi, et autres
Publié: (2026)
When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints
par: Chen, Yuheng, et autres
Publié: (2026)
par: Chen, Yuheng, et autres
Publié: (2026)
SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?
par: Chen, Jiamin, et autres
Publié: (2026)
par: Chen, Jiamin, et autres
Publié: (2026)
How Reliable is Multilingual LLM-as-a-Judge?
par: Fu, Xiyan, et autres
Publié: (2025)
par: Fu, Xiyan, et autres
Publié: (2025)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
par: Chen, Dongping, et autres
Publié: (2024)
par: Chen, Dongping, et autres
Publié: (2024)
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
par: Vishnubhotla, Krishnapriya, et autres
Publié: (2026)
par: Vishnubhotla, Krishnapriya, et autres
Publié: (2026)
UBench: Benchmarking Uncertainty in Large Language Models with Multiple Choice Questions
par: Wang, Xunzhi, et autres
Publié: (2024)
par: Wang, Xunzhi, et autres
Publié: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
par: Tan, Sijun, et autres
Publié: (2024)
par: Tan, Sijun, et autres
Publié: (2024)
SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
par: Zhu, Xiangyang, et autres
Publié: (2025)
par: Zhu, Xiangyang, et autres
Publié: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
par: Zhou, Yilun, et autres
Publié: (2025)
par: Zhou, Yilun, et autres
Publié: (2025)
Improving Score Reliability of Multiple Choice Benchmarks with Consistency Evaluation and Altered Answer Choices
par: Cavalin, Paulo, et autres
Publié: (2025)
par: Cavalin, Paulo, et autres
Publié: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
par: Jiang, Hongchao, et autres
Publié: (2025)
par: Jiang, Hongchao, et autres
Publié: (2025)
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
par: Zhou, Yuhang, et autres
Publié: (2025)
par: Zhou, Yuhang, et autres
Publié: (2025)
Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation
par: Sternlicht, Noy, et autres
Publié: (2025)
par: Sternlicht, Noy, et autres
Publié: (2025)
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation
par: Hidayat, Naila Shafirni, et autres
Publié: (2025)
par: Hidayat, Naila Shafirni, et autres
Publié: (2025)
MM-JudgeBias: A Benchmark for Evaluating Compositional Biases in MLLM-as-a-Judge
par: Lee, Sua, et autres
Publié: (2026)
par: Lee, Sua, et autres
Publié: (2026)
PERSPECTRA: A Scalable and Configurable Pluralist Benchmark of Perspectives from Arguments
par: Nie, Shangrui, et autres
Publié: (2026)
par: Nie, Shangrui, et autres
Publié: (2026)
Configurable Safety Tuning of Language Models with Synthetic Preference Data
par: Gallego, Victor
Publié: (2024)
par: Gallego, Victor
Publié: (2024)
CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems
par: Yang, Jingbo, et autres
Publié: (2026)
par: Yang, Jingbo, et autres
Publié: (2026)
How are Prompts Different in Terms of Sensitivity?
par: Lu, Sheng, et autres
Publié: (2023)
par: Lu, Sheng, et autres
Publié: (2023)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
par: Yang, Bo, et autres
Publié: (2026)
par: Yang, Bo, et autres
Publié: (2026)
Attribution Quality in AI-Generated Content:Benchmarking Style Embeddings and LLM Judges
par: Abbas, Misam
Publié: (2025)
par: Abbas, Misam
Publié: (2025)
The Latin Substrate: How Language Models Represent and Mediate Script Choice
par: Gurgurov, Daniil, et autres
Publié: (2026)
par: Gurgurov, Daniil, et autres
Publié: (2026)
BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
par: Balepur, Nishant, et autres
Publié: (2026)
par: Balepur, Nishant, et autres
Publié: (2026)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
par: Cai, Yunna, et autres
Publié: (2025)
par: Cai, Yunna, et autres
Publié: (2025)
Automated Safety Benchmarking: A Multi-agent Pipeline for LVLMs
par: Zhu, Xiangyang, et autres
Publié: (2026)
par: Zhu, Xiangyang, et autres
Publié: (2026)
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
par: Zhang, Yushun, et autres
Publié: (2026)
par: Zhang, Yushun, et autres
Publié: (2026)
MLLM-as-a-Judge for Image Safety without Human Labeling
par: Wang, Zhenting, et autres
Publié: (2024)
par: Wang, Zhenting, et autres
Publié: (2024)
A Judge-free LLM Open-ended Generation Benchmark Based on the Distributional Hypothesis
par: Imajo, Kentaro, et autres
Publié: (2025)
par: Imajo, Kentaro, et autres
Publié: (2025)
MR. Judge: Multimodal Reasoner as a Judge
par: Pi, Renjie, et autres
Publié: (2025)
par: Pi, Renjie, et autres
Publié: (2025)
Documents similaires
-
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
par: Bellibatlu, Rohith Reddy, et autres
Publié: (2026) -
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation
par: Zhang, Xinran
Publié: (2026) -
An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability
par: Yamauchi, Yusuke, et autres
Publié: (2025) -
Beyond Creed: A Non-Identity Safety Condition A Strong Empirical Alternative to Identity Framing in Low-Data LoRA Fine-Tuning
par: Zhang, Xinran
Publié: (2026) -
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
par: Yuan, Tongxin, et autres
Publié: (2024)