Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Collot, Stephane, Fraser, Colin, Zhao, Justin, Shen, William F., Willi, Timon, Leontiadis, Ilias |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
von: Shen, William F., et al.
Veröffentlicht: (2026)
von: Shen, William F., et al.
Veröffentlicht: (2026)
Training AI Co-Scientists Using Rubric Rewards
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
von: Goel, Shashwat, et al.
Veröffentlicht: (2025)
Evaluating Privacy Leakage in Split Learning
von: Qiu, Xinchi, et al.
Veröffentlicht: (2023)
von: Qiu, Xinchi, et al.
Veröffentlicht: (2023)
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
von: Clegg, Kester, et al.
Veröffentlicht: (2025)
Youden's Demon is Sylvester's Problem
von: Frick, Florian, et al.
Veröffentlicht: (2024)
von: Frick, Florian, et al.
Veröffentlicht: (2024)
Youden's demon is Sylvester's problem
von: Florian Frick, et al.
Veröffentlicht: (2025)
von: Florian Frick, et al.
Veröffentlicht: (2025)
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
von: Lupu, Andrei, et al.
Veröffentlicht: (2025)
von: Lupu, Andrei, et al.
Veröffentlicht: (2025)
Biomarkers selection and combination based on the weighted Youden index
von: Sun, Ao, et al.
Veröffentlicht: (2025)
von: Sun, Ao, et al.
Veröffentlicht: (2025)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
Causal Judge Evaluation: Calibrated Surrogate Metrics for LLM Systems
von: Landesberg, Eddie, et al.
Veröffentlicht: (2025)
von: Landesberg, Eddie, et al.
Veröffentlicht: (2025)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
Evaluating LLM Metrics Through Real-World Capabilities
von: Miller, Justin K, et al.
Veröffentlicht: (2025)
von: Miller, Justin K, et al.
Veröffentlicht: (2025)
LLM as a Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation
von: Eigler, Lukáš, et al.
Veröffentlicht: (2026)
von: Eigler, Lukáš, et al.
Veröffentlicht: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
LLM-as-a-Judge for Scalable Test Coverage Evaluation: Accuracy, Operational Reliability, and Cost
von: Huang, Donghao, et al.
Veröffentlicht: (2025)
von: Huang, Donghao, et al.
Veröffentlicht: (2025)
AcceLLM: Accelerating LLM Inference using Redundancy for Load Balancing and Data Locality
von: Bournias, Ilias, et al.
Veröffentlicht: (2024)
von: Bournias, Ilias, et al.
Veröffentlicht: (2024)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
Biomarker combination based on the Youden index with and without gold standard
von: Sun, Ao, et al.
Veröffentlicht: (2024)
von: Sun, Ao, et al.
Veröffentlicht: (2024)
Optimal Linear Combination of Biomarkers by Weighted Youden Index Maximization
von: Sizhe Wang, et al.
Veröffentlicht: (2025)
von: Sizhe Wang, et al.
Veröffentlicht: (2025)
Estimation of Multi‐Category Youden Index Based on the Lehmann Assumption
von: Qunqiang Feng, et al.
Veröffentlicht: (2025)
von: Qunqiang Feng, et al.
Veröffentlicht: (2025)
Are We on the Right Way to Assessing LLM-as-a-Judge?
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
von: Feng, Yuanning, et al.
Veröffentlicht: (2025)
Multi-Agent LLM Judge: automatic personalized LLM judge design for evaluating natural language generation applications
von: Cao, Hongliu, et al.
Veröffentlicht: (2025)
von: Cao, Hongliu, et al.
Veröffentlicht: (2025)
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates
von: Wei, Hui, et al.
Veröffentlicht: (2024)
von: Wei, Hui, et al.
Veröffentlicht: (2024)
Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards
von: Su, Jinyan, et al.
Veröffentlicht: (2025)
von: Su, Jinyan, et al.
Veröffentlicht: (2025)
Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation
von: Cao, Hongliu, et al.
Veröffentlicht: (2026)
von: Cao, Hongliu, et al.
Veröffentlicht: (2026)
Semiparametric Joint Inference for Sensitivity and Specificity at the Youden-Optimal Cut-Off
von: Liu, Siyan, et al.
Veröffentlicht: (2026)
von: Liu, Siyan, et al.
Veröffentlicht: (2026)
Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy
von: Deviyani, Athiya, et al.
Veröffentlicht: (2025)
von: Deviyani, Athiya, et al.
Veröffentlicht: (2025)
When Metrics Disagree: Automatic Similarity vs. LLM-as-a-Judge for Clinical Dialogue Evaluation
von: Sun, Bian, et al.
Veröffentlicht: (2026)
von: Sun, Bian, et al.
Veröffentlicht: (2026)
Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
von: Chalkidis, Ilias
Veröffentlicht: (2025)
von: Chalkidis, Ilias
Veröffentlicht: (2025)
Judging for ourselves
von: Justin Khoo
Veröffentlicht: (2024)
von: Justin Khoo
Veröffentlicht: (2024)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
Poisson and Gaussian approximations of the power divergence family of statistics
von: Daly, Fraser
Veröffentlicht: (2023)
von: Daly, Fraser
Veröffentlicht: (2023)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
von: Li, Xiaochuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaochuan, et al.
Veröffentlicht: (2025)
Methodological Approaches for the Estimation of Confidence Intervals on Partial Youden Index Under Verification Bias
von: Sihan Jia, et al.
Veröffentlicht: (2026)
von: Sihan Jia, et al.
Veröffentlicht: (2026)
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
von: Weng, Shihao, et al.
Veröffentlicht: (2026)
von: Weng, Shihao, et al.
Veröffentlicht: (2026)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
von: Zhu, Ziyi, et al.
Veröffentlicht: (2026)
von: Zhu, Ziyi, et al.
Veröffentlicht: (2026)
Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges
von: Eiras, Francisco, et al.
Veröffentlicht: (2025)
von: Eiras, Francisco, et al.
Veröffentlicht: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
von: Tong, Terry, et al.
Veröffentlicht: (2025)
von: Tong, Terry, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks
von: Shen, William F., et al.
Veröffentlicht: (2026) -
Training AI Co-Scientists Using Rubric Rewards
von: Goel, Shashwat, et al.
Veröffentlicht: (2025) -
Evaluating Privacy Leakage in Split Learning
von: Qiu, Xinchi, et al.
Veröffentlicht: (2023) -
Evaluating Metrics for Safety with LLM-as-Judges
von: Clegg, Kester, et al.
Veröffentlicht: (2025) -
Youden's Demon is Sylvester's Problem
von: Frick, Florian, et al.
Veröffentlicht: (2024)