CLEAR: A Clinically-Grounded Tabular Framework for Radiology Report Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918143978373120 |
|---|---|
| author | Jiang, Yuyang Chen, Chacha Wang, Shengyuan Li, Feng Tang, Zecong Mervak, Benjamin M. Chelala, Lydia Straus, Christopher M Chahine, Reve Armato III, Samuel G. Tan, Chenhao |
| author_facet | Jiang, Yuyang Chen, Chacha Wang, Shengyuan Li, Feng Tang, Zecong Mervak, Benjamin M. Chelala, Lydia Straus, Christopher M Chahine, Reve Armato III, Samuel G. Tan, Chenhao |
| contents | Existing metrics often lack the granularity and interpretability to capture nuanced clinical differences between candidate and ground-truth radiology reports, resulting in suboptimal evaluation. We introduce a Clinically-grounded tabular framework with Expert-curated labels and Attribute-level comparison for Radiology report evaluation (CLEAR). CLEAR not only examines whether a report can accurately identify the presence or absence of medical conditions, but also assesses whether it can precisely describe each positively identified condition across five key attributes: first occurrence, change, severity, descriptive location, and recommendation. Compared to prior works, CLEAR's multi-dimensional, attribute-level outputs enable a more comprehensive and clinically interpretable evaluation of report quality. Additionally, to measure the clinical alignment of CLEAR, we collaborate with five board-certified radiologists to develop CLEAR-Bench, a dataset of 100 chest X-ray reports from MIMIC-CXR, annotated across 6 curated attributes and 13 CheXpert conditions. Our experiments show that CLEAR achieves high accuracy in extracting clinical attributes and provides automated metrics that are strongly aligned with clinical judgment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_16325 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | CLEAR: A Clinically-Grounded Tabular Framework for Radiology Report Evaluation Jiang, Yuyang Chen, Chacha Wang, Shengyuan Li, Feng Tang, Zecong Mervak, Benjamin M. Chelala, Lydia Straus, Christopher M Chahine, Reve Armato III, Samuel G. Tan, Chenhao Computation and Language Artificial Intelligence Computers and Society Existing metrics often lack the granularity and interpretability to capture nuanced clinical differences between candidate and ground-truth radiology reports, resulting in suboptimal evaluation. We introduce a Clinically-grounded tabular framework with Expert-curated labels and Attribute-level comparison for Radiology report evaluation (CLEAR). CLEAR not only examines whether a report can accurately identify the presence or absence of medical conditions, but also assesses whether it can precisely describe each positively identified condition across five key attributes: first occurrence, change, severity, descriptive location, and recommendation. Compared to prior works, CLEAR's multi-dimensional, attribute-level outputs enable a more comprehensive and clinically interpretable evaluation of report quality. Additionally, to measure the clinical alignment of CLEAR, we collaborate with five board-certified radiologists to develop CLEAR-Bench, a dataset of 100 chest X-ray reports from MIMIC-CXR, annotated across 6 curated attributes and 13 CheXpert conditions. Our experiments show that CLEAR achieves high accuracy in extracting clinical attributes and provides automated metrics that are strongly aligned with clinical judgment. |
| title | CLEAR: A Clinically-Grounded Tabular Framework for Radiology Report Evaluation |
| topic | Computation and Language Artificial Intelligence Computers and Society |
| url | https://arxiv.org/abs/2505.16325 |