VERT: Reliable LLM Judges for Radiology Report Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Bologna, Federica, Corbeil, Jean-Philippe, Wilkens, Matthew, Abacha, Asma Ben |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CQA-Eval: Designing Reliable Evaluations of Multi-paragraph Clinical QA under Resource Constraints
by: Bologna, Federica, et al.
Published: (2025)
by: Bologna, Federica, et al.
Published: (2025)
Overview of the MEDIQA-OE 2025 Shared Task on Medical Order Extraction from Doctor-Patient Consultations
by: Corbeil, Jean-Philippe, et al.
Published: (2025)
by: Corbeil, Jean-Philippe, et al.
Published: (2025)
A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment
by: Corbeil, Jean-Philippe, et al.
Published: (2025)
by: Corbeil, Jean-Philippe, et al.
Published: (2025)
IryoNLP at MEDIQA-CORR 2024: Tackling the Medical Error Detection & Correction Task On the Shoulders of Medical Agents
by: Corbeil, Jean-Philippe
Published: (2024)
by: Corbeil, Jean-Philippe
Published: (2024)
Empowering Healthcare Practitioners with Language Models: Structuring Speech Transcripts in Two Real-World Clinical Applications
by: Corbeil, Jean-Philippe, et al.
Published: (2025)
by: Corbeil, Jean-Philippe, et al.
Published: (2025)
GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents
by: Moll, Johannes, et al.
Published: (2026)
by: Moll, Johannes, et al.
Published: (2026)
MRScore: Evaluating Radiology Report Generation with LLM-based Reward System
by: Liu, Yunyi, et al.
Published: (2024)
by: Liu, Yunyi, et al.
Published: (2024)
GREEN: Generative Radiology Report Evaluation and Error Notation
by: Ostmeier, Sophie, et al.
Published: (2024)
by: Ostmeier, Sophie, et al.
Published: (2024)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
by: Hong, Yihan, et al.
Published: (2026)
by: Hong, Yihan, et al.
Published: (2026)
Coarse-to-Fine Personalized LLM Impressions for Streamlined Radiology Reports
by: Sun, Chengbo, et al.
Published: (2025)
by: Sun, Chengbo, et al.
Published: (2025)
Evaluating Metrics for Safety with LLM-as-Judges
by: Clegg, Kester, et al.
Published: (2025)
by: Clegg, Kester, et al.
Published: (2025)
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
by: Abacha, Asma Ben, et al.
Published: (2024)
by: Abacha, Asma Ben, et al.
Published: (2024)
PARROT: An Open Multilingual Radiology Reports Dataset
by: Guellec, Bastien Le, et al.
Published: (2025)
by: Guellec, Bastien Le, et al.
Published: (2025)
LLM-RadJudge: Achieving Radiologist-Level Evaluation for X-Ray Report Generation
by: Wang, Zilong, et al.
Published: (2024)
by: Wang, Zilong, et al.
Published: (2024)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
by: Schwinn, Leo, et al.
Published: (2026)
by: Schwinn, Leo, et al.
Published: (2026)
RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores
by: Li, Yingshu, et al.
Published: (2025)
by: Li, Yingshu, et al.
Published: (2025)
Leveraging Professional Radiologists' Expertise to Enhance LLMs' Evaluation for Radiology Reports
by: Zhu, Qingqing, et al.
Published: (2024)
by: Zhu, Qingqing, et al.
Published: (2024)
DermaVQA-DAS: Dermatology Assessment Schema (DAS) & Datasets for Closed-Ended Question Answering & Segmentation in Patient-Generated Dermatology Images
by: Yim, Wen-wai, et al.
Published: (2025)
by: Yim, Wen-wai, et al.
Published: (2025)
RadAnnotate: Large Language Models for Efficient and Reliable Radiology Report Annotation
by: Shetty, Saisha Pradeep, et al.
Published: (2026)
by: Shetty, Saisha Pradeep, et al.
Published: (2026)
Process Reward Models for Sentence-Level Verification of LVLM Radiology Reports
by: Thomas, Alois, et al.
Published: (2025)
by: Thomas, Alois, et al.
Published: (2025)
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
by: Baharoon, Mohammed, et al.
Published: (2026)
by: Baharoon, Mohammed, et al.
Published: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
Standardizing Longitudinal Radiology Report Evaluation via Large Language Model Annotation
by: Wang, Xinyi, et al.
Published: (2026)
by: Wang, Xinyi, et al.
Published: (2026)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
by: Thakur, Aman Singh, et al.
Published: (2024)
by: Thakur, Aman Singh, et al.
Published: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
by: Saha, Swarnadeep, et al.
Published: (2025)
by: Saha, Swarnadeep, et al.
Published: (2025)
Summarizing Radiology Reports Findings into Impressions
by: de Padua, Raul Salles, et al.
Published: (2024)
by: de Padua, Raul Salles, et al.
Published: (2024)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
by: Gupta, Manan, et al.
Published: (2026)
by: Gupta, Manan, et al.
Published: (2026)
CLEAR: A Clinically-Grounded Tabular Framework for Radiology Report Evaluation
by: Jiang, Yuyang, et al.
Published: (2025)
by: Jiang, Yuyang, et al.
Published: (2025)
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge
by: Shi, Lin, et al.
Published: (2024)
by: Shi, Lin, et al.
Published: (2024)
LeMAJ (Legal LLM-as-a-Judge): Bridging Legal Reasoning and LLM Evaluation
by: Enguehard, Joseph, et al.
Published: (2025)
by: Enguehard, Joseph, et al.
Published: (2025)
Two-Pronged Human Evaluation of ChatGPT Self-Correction in Radiology Report Simplification
by: Yang, Ziyu, et al.
Published: (2024)
by: Yang, Ziyu, et al.
Published: (2024)
Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Raters
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation
by: Liu, Yunyi, et al.
Published: (2024)
by: Liu, Yunyi, et al.
Published: (2024)
Do Repetitions Matter? Strengthening Reliability in LLM Evaluations
by: Gonzalez, Miguel Angel Alvarado, et al.
Published: (2025)
by: Gonzalez, Miguel Angel Alvarado, et al.
Published: (2025)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
by: Jin, Jiho, et al.
Published: (2026)
by: Jin, Jiho, et al.
Published: (2026)
Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems
by: Chhabra, Mukul, et al.
Published: (2026)
by: Chhabra, Mukul, et al.
Published: (2026)
Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
by: Verga, Pat, et al.
Published: (2024)
by: Verga, Pat, et al.
Published: (2024)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
by: Shi, Zhichao, et al.
Published: (2025)
by: Shi, Zhichao, et al.
Published: (2025)
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
by: Wang, Yidong, et al.
Published: (2025)
by: Wang, Yidong, et al.
Published: (2025)
Similar Items
-
CQA-Eval: Designing Reliable Evaluations of Multi-paragraph Clinical QA under Resource Constraints
by: Bologna, Federica, et al.
Published: (2025) -
Overview of the MEDIQA-OE 2025 Shared Task on Medical Order Extraction from Doctor-Patient Consultations
by: Corbeil, Jean-Philippe, et al.
Published: (2025) -
A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment
by: Corbeil, Jean-Philippe, et al.
Published: (2025) -
IryoNLP at MEDIQA-CORR 2024: Tackling the Medical Error Detection & Correction Task On the Shoulders of Medical Agents
by: Corbeil, Jean-Philippe
Published: (2024) -
Empowering Healthcare Practitioners with Language Models: Structuring Speech Transcripts in Two Real-World Clinical Applications
by: Corbeil, Jean-Philippe, et al.
Published: (2025)