Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
Fuente:
arXiv
Saved in:
| Main Authors: | Collot, Stephane, Fraser, Colin, Zhao, Justin, Shen, William F., Willi, Timon, Leontiadis, Ilias |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
by: Shen, William F., et al.
Published: (2026)
by: Shen, William F., et al.
Published: (2026)
Training AI Co-Scientists Using Rubric Rewards
by: Goel, Shashwat, et al.
Published: (2025)
by: Goel, Shashwat, et al.
Published: (2025)
Evaluating Privacy Leakage in Split Learning
by: Qiu, Xinchi, et al.
Published: (2023)
by: Qiu, Xinchi, et al.
Published: (2023)
Evaluating Metrics for Safety with LLM-as-Judges
by: Clegg, Kester, et al.
Published: (2025)
by: Clegg, Kester, et al.
Published: (2025)
Youden's Demon is Sylvester's Problem
by: Frick, Florian, et al.
Published: (2024)
by: Frick, Florian, et al.
Published: (2024)
Youden's demon is Sylvester's problem
by: Florian Frick, et al.
Published: (2025)
by: Florian Frick, et al.
Published: (2025)
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
by: Lupu, Andrei, et al.
Published: (2025)
by: Lupu, Andrei, et al.
Published: (2025)
Biomarkers selection and combination based on the weighted Youden index
by: Sun, Ao, et al.
Published: (2025)
by: Sun, Ao, et al.
Published: (2025)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Causal Judge Evaluation: Calibrated Surrogate Metrics for LLM Systems
by: Landesberg, Eddie, et al.
Published: (2025)
by: Landesberg, Eddie, et al.
Published: (2025)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Evaluating LLM Metrics Through Real-World Capabilities
by: Miller, Justin K, et al.
Published: (2025)
by: Miller, Justin K, et al.
Published: (2025)
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
by: Eigler, Lukáš, et al.
Published: (2026)
by: Eigler, Lukáš, et al.
Published: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
by: Alam, Firoj, et al.
Published: (2026)
by: Alam, Firoj, et al.
Published: (2026)
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
by: Huang, Donghao, et al.
Published: (2025)
by: Huang, Donghao, et al.
Published: (2025)
AcceLLM: Accelerating LLM Inference using Redundancy for Load Balancing and Data Locality
by: Bournias, Ilias, et al.
Published: (2024)
by: Bournias, Ilias, et al.
Published: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
Biomarker combination based on the Youden index with and without gold standard
by: Sun, Ao, et al.
Published: (2024)
by: Sun, Ao, et al.
Published: (2024)
Optimal Linear Combination of Biomarkers by Weighted Youden Index Maximization
by: Sizhe Wang, et al.
Published: (2025)
by: Sizhe Wang, et al.
Published: (2025)
Estimation of Multi‐Category Youden Index Based on the Lehmann Assumption
by: Qunqiang Feng, et al.
Published: (2025)
by: Qunqiang Feng, et al.
Published: (2025)
Are We on the Right Way to Assessing LLM-as-a-Judge?
by: Feng, Yuanning, et al.
Published: (2025)
by: Feng, Yuanning, et al.
Published: (2025)
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
by: Cao, Hongliu, et al.
Published: (2025)
by: Cao, Hongliu, et al.
Published: (2025)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
by: Wei, Hui, et al.
Published: (2024)
by: Wei, Hui, et al.
Published: (2024)
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation
by: Cao, Hongliu, et al.
Published: (2026)
by: Cao, Hongliu, et al.
Published: (2026)
Semiparametric Joint Inference for Sensitivity and Specificity at the Youden-Optimal Cut-Off
by: Liu, Siyan, et al.
Published: (2026)
by: Liu, Siyan, et al.
Published: (2026)
Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy
by: Deviyani, Athiya, et al.
Published: (2025)
by: Deviyani, Athiya, et al.
Published: (2025)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
by: Sun, Bian, et al.
Published: (2026)
by: Sun, Bian, et al.
Published: (2026)
Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
by: Chalkidis, Ilias
Published: (2025)
by: Chalkidis, Ilias
Published: (2025)
Judging for ourselves
by: Justin Khoo
Published: (2024)
by: Justin Khoo
Published: (2024)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
by: Zhou, Yilun, et al.
Published: (2025)
by: Zhou, Yilun, et al.
Published: (2025)
Poisson and Gaussian approximations of the power divergence family of statistics
by: Daly, Fraser
Published: (2023)
by: Daly, Fraser
Published: (2023)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
by: Li, Xiaochuan, et al.
Published: (2025)
by: Li, Xiaochuan, et al.
Published: (2025)
Methodological Approaches for the Estimation of Confidence Intervals on Partial Youden Index Under Verification Bias
by: Sihan Jia, et al.
Published: (2026)
by: Sihan Jia, et al.
Published: (2026)
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
by: Weng, Shihao, et al.
Published: (2026)
by: Weng, Shihao, et al.
Published: (2026)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
by: Zhu, Ziyi, et al.
Published: (2026)
by: Zhu, Ziyi, et al.
Published: (2026)
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
by: Eiras, Francisco, et al.
Published: (2025)
by: Eiras, Francisco, et al.
Published: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
by: Tong, Terry, et al.
Published: (2025)
by: Tong, Terry, et al.
Published: (2025)
Similar Items
-
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
by: Shen, William F., et al.
Published: (2026) -
Training AI Co-Scientists Using Rubric Rewards
by: Goel, Shashwat, et al.
Published: (2025) -
Evaluating Privacy Leakage in Split Learning
by: Qiu, Xinchi, et al.
Published: (2023) -
Evaluating Metrics for Safety with LLM-as-Judges
by: Clegg, Kester, et al.
Published: (2025) -
Youden's Demon is Sylvester's Problem
by: Frick, Florian, et al.
Published: (2024)