Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
Fuente:
arXiv
Salvato in:
| Autori principali: | Belmadani, Ikram, Khettari, Oumaima El, Beaufils, Pacôme Constant dit, Dufour, Richard, Favre, Benoit |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MedInjection-FR: Exploring the Role of Native, Synthetic, and Translated Data in Biomedical Instruction Tuning
di: Belmadani, Ikram, et al.
Pubblicazione: (2026)
di: Belmadani, Ikram, et al.
Pubblicazione: (2026)
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain
di: Labrak, Yanis, et al.
Pubblicazione: (2024)
di: Labrak, Yanis, et al.
Pubblicazione: (2024)
LLM, Reporting In! Medical Information Extraction Across Prompting, Fine-tuning and Post-correction
di: Belmadani, Ikram, et al.
Pubblicazione: (2025)
di: Belmadani, Ikram, et al.
Pubblicazione: (2025)
Improving Social Determinants of Health Documentation in French EHRs Using Large Language Models
di: Bazoge, Adrien, et al.
Pubblicazione: (2025)
di: Bazoge, Adrien, et al.
Pubblicazione: (2025)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
di: Lee, Dongryeol, et al.
Pubblicazione: (2026)
di: Lee, Dongryeol, et al.
Pubblicazione: (2026)
Brain abscess and heart: the phantom menace?
di: Pacôme Constant dit Beaufils, et al.
Pubblicazione: (2024)
di: Pacôme Constant dit Beaufils, et al.
Pubblicazione: (2024)
Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation
di: Adib, Shefayat E Shams, et al.
Pubblicazione: (2026)
di: Adib, Shefayat E Shams, et al.
Pubblicazione: (2026)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
di: Thakur, Aman Singh, et al.
Pubblicazione: (2024)
di: Thakur, Aman Singh, et al.
Pubblicazione: (2024)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
di: Zhu, Ziyi, et al.
Pubblicazione: (2026)
di: Zhu, Ziyi, et al.
Pubblicazione: (2026)
The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
di: Marioriyad, Arash, et al.
Pubblicazione: (2026)
di: Marioriyad, Arash, et al.
Pubblicazione: (2026)
Evaluating Metrics for Safety with LLM-as-Judges
di: Clegg, Kester, et al.
Pubblicazione: (2025)
di: Clegg, Kester, et al.
Pubblicazione: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024)
di: Tan, Sijun, et al.
Pubblicazione: (2024)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
di: Shi, Lin, et al.
Pubblicazione: (2024)
di: Shi, Lin, et al.
Pubblicazione: (2024)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
Summarization for Generative Relation Extraction in the Microbiome Domain
di: Khettari, Oumaima El, et al.
Pubblicazione: (2025)
di: Khettari, Oumaima El, et al.
Pubblicazione: (2025)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
di: Yang, Bo, et al.
Pubblicazione: (2026)
di: Yang, Bo, et al.
Pubblicazione: (2026)
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
di: Marioriyad, Arash, et al.
Pubblicazione: (2025)
di: Marioriyad, Arash, et al.
Pubblicazione: (2025)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
di: Moon, Jiwon, et al.
Pubblicazione: (2025)
di: Moon, Jiwon, et al.
Pubblicazione: (2025)
Reassessing Extractive QA Datasets at Scale: LLM-as-a-Judge and In-Depth Analyses
di: Ho, Xanh, et al.
Pubblicazione: (2025)
di: Ho, Xanh, et al.
Pubblicazione: (2025)
BadJudge: Backdoor Vulnerabilities of LLM-as-a-Judge
di: Tong, Terry, et al.
Pubblicazione: (2025)
di: Tong, Terry, et al.
Pubblicazione: (2025)
Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
di: Liu, Shuliang, et al.
Pubblicazione: (2025)
di: Liu, Shuliang, et al.
Pubblicazione: (2025)
Evaluating Scoring Bias in LLM-as-a-Judge
di: Li, Qingquan, et al.
Pubblicazione: (2025)
di: Li, Qingquan, et al.
Pubblicazione: (2025)
Quantitative LLM Judges
di: Sahoo, Aishwarya, et al.
Pubblicazione: (2025)
di: Sahoo, Aishwarya, et al.
Pubblicazione: (2025)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
di: Bellibatlu, Rohith Reddy, et al.
Pubblicazione: (2026)
di: Bellibatlu, Rohith Reddy, et al.
Pubblicazione: (2026)
Rethinking Atomic Decomposition for LLM Judges: A Prompt-Controlled Study of Reference-Grounded QA Evaluation
di: Zhang, Xinran
Pubblicazione: (2026)
di: Zhang, Xinran
Pubblicazione: (2026)
MR. Judge: Multimodal Reasoner as a Judge
di: Pi, Renjie, et al.
Pubblicazione: (2025)
di: Pi, Renjie, et al.
Pubblicazione: (2025)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
di: Wang, Yidong, et al.
Pubblicazione: (2025)
di: Wang, Yidong, et al.
Pubblicazione: (2025)
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
di: Huang, Hui, et al.
Pubblicazione: (2024)
di: Huang, Hui, et al.
Pubblicazione: (2024)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
di: Tang, Zhenwei, et al.
Pubblicazione: (2026)
di: Tang, Zhenwei, et al.
Pubblicazione: (2026)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
di: Liu, Yixin, et al.
Pubblicazione: (2025)
di: Liu, Yixin, et al.
Pubblicazione: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
di: Jiang, Hongchao, et al.
Pubblicazione: (2025)
di: Jiang, Hongchao, et al.
Pubblicazione: (2025)
Familial Reversible Cerebral Vasoconstriction Syndrome: Insights From Two Families
di: Pacôme Constant dit Beaufils, et al.
Pubblicazione: (2025)
di: Pacôme Constant dit Beaufils, et al.
Pubblicazione: (2025)
Can LLM be a Personalized Judge?
di: Dong, Yijiang River, et al.
Pubblicazione: (2024)
di: Dong, Yijiang River, et al.
Pubblicazione: (2024)
A Judge-free LLM Open-ended Generation Benchmark Based on the Distributional Hypothesis
di: Imajo, Kentaro, et al.
Pubblicazione: (2025)
di: Imajo, Kentaro, et al.
Pubblicazione: (2025)
Judge's Verdict: A Comprehensive Analysis of LLM Judge Capability Through Human Agreement
di: Han, Steve, et al.
Pubblicazione: (2025)
di: Han, Steve, et al.
Pubblicazione: (2025)
VERT: Reliable LLM Judges for Radiology Report Evaluation
di: Bologna, Federica, et al.
Pubblicazione: (2026)
di: Bologna, Federica, et al.
Pubblicazione: (2026)
Cross-Lingual LLM-Judge Transfer via Evaluation Decomposition
di: Sheth, Ivaxi, et al.
Pubblicazione: (2026)
di: Sheth, Ivaxi, et al.
Pubblicazione: (2026)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
di: Chen, Junjie, et al.
Pubblicazione: (2026)
di: Chen, Junjie, et al.
Pubblicazione: (2026)
JudgeLRM: Large Reasoning Models as a Judge
di: Chen, Nuo, et al.
Pubblicazione: (2025)
di: Chen, Nuo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
MedInjection-FR: Exploring the Role of Native, Synthetic, and Translated Data in Biomedical Instruction Tuning
di: Belmadani, Ikram, et al.
Pubblicazione: (2026) -
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain
di: Labrak, Yanis, et al.
Pubblicazione: (2024) -
LLM, Reporting In! Medical Information Extraction Across Prompting, Fine-tuning and Post-correction
di: Belmadani, Ikram, et al.
Pubblicazione: (2025) -
Improving Social Determinants of Health Documentation in French EHRs Using Large Language Models
di: Bazoge, Adrien, et al.
Pubblicazione: (2025) -
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
di: Lee, Dongryeol, et al.
Pubblicazione: (2026)