FineRadScore: A Radiology Report Line-by-Line Evaluation Technique Generating Corrections with Severity Scores

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Alyssa, Banerjee, Oishi, Wu, Kay, Reis, Eduardo Pontes, Rajpurkar, Pranav
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914908318203904
author Huang, Alyssa
Banerjee, Oishi
Wu, Kay
Reis, Eduardo Pontes
Rajpurkar, Pranav
author_facet Huang, Alyssa
Banerjee, Oishi
Wu, Kay
Reis, Eduardo Pontes
Rajpurkar, Pranav
contents The current gold standard for evaluating generated chest x-ray (CXR) reports is through radiologist annotations. However, this process can be extremely time-consuming and costly, especially when evaluating large numbers of reports. In this work, we present FineRadScore, a Large Language Model (LLM)-based automated evaluation metric for generated CXR reports. Given a candidate report and a ground-truth report, FineRadScore gives the minimum number of line-by-line corrections required to go from the candidate to the ground-truth report. Additionally, FineRadScore provides an error severity rating with each correction and generates comments explaining why the correction was needed. We demonstrate that FineRadScore's corrections and error severity scores align with radiologist opinions. We also show that, when used to judge the quality of the report as a whole, FineRadScore aligns with radiologists as well as current state-of-the-art automated CXR evaluation metrics. Finally, we analyze FineRadScore's shortcomings to provide suggestions for future improvements.
format Preprint
id arxiv_https___arxiv_org_abs_2405_20613
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FineRadScore: A Radiology Report Line-by-Line Evaluation Technique Generating Corrections with Severity Scores
Huang, Alyssa
Banerjee, Oishi
Wu, Kay
Reis, Eduardo Pontes
Rajpurkar, Pranav
Computation and Language
The current gold standard for evaluating generated chest x-ray (CXR) reports is through radiologist annotations. However, this process can be extremely time-consuming and costly, especially when evaluating large numbers of reports. In this work, we present FineRadScore, a Large Language Model (LLM)-based automated evaluation metric for generated CXR reports. Given a candidate report and a ground-truth report, FineRadScore gives the minimum number of line-by-line corrections required to go from the candidate to the ground-truth report. Additionally, FineRadScore provides an error severity rating with each correction and generates comments explaining why the correction was needed. We demonstrate that FineRadScore's corrections and error severity scores align with radiologist opinions. We also show that, when used to judge the quality of the report as a whole, FineRadScore aligns with radiologists as well as current state-of-the-art automated CXR evaluation metrics. Finally, we analyze FineRadScore's shortcomings to provide suggestions for future improvements.
title FineRadScore: A Radiology Report Line-by-Line Evaluation Technique Generating Corrections with Severity Scores
topic Computation and Language
url https://arxiv.org/abs/2405.20613