Leveraging Professional Radiologists' Expertise to Enhance LLMs' Evaluation for Radiology Reports

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhu, Qingqing, Chen, Xiuying, Jin, Qiao, Hou, Benjamin, Mathai, Tejas Sudharshan, Mukherjee, Pritam, Gao, Xin, Summers, Ronald M, Lu, Zhiyong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911778703671296
author Zhu, Qingqing
Chen, Xiuying
Jin, Qiao
Hou, Benjamin
Mathai, Tejas Sudharshan
Mukherjee, Pritam
Gao, Xin
Summers, Ronald M
Lu, Zhiyong
author_facet Zhu, Qingqing
Chen, Xiuying
Jin, Qiao
Hou, Benjamin
Mathai, Tejas Sudharshan
Mukherjee, Pritam
Gao, Xin
Summers, Ronald M
Lu, Zhiyong
contents In radiology, Artificial Intelligence (AI) has significantly advanced report generation, but automatic evaluation of these AI-produced reports remains challenging. Current metrics, such as Conventional Natural Language Generation (NLG) and Clinical Efficacy (CE), often fall short in capturing the semantic intricacies of clinical contexts or overemphasize clinical details, undermining report clarity. To overcome these issues, our proposed method synergizes the expertise of professional radiologists with Large Language Models (LLMs), like GPT-3.5 and GPT-4 1. Utilizing In-Context Instruction Learning (ICIL) and Chain of Thought (CoT) reasoning, our approach aligns LLM evaluations with radiologist standards, enabling detailed comparisons between human and AI generated reports. This is further enhanced by a Regression model that aggregates sentence evaluation scores. Experimental results show that our "Detailed GPT-4 (5-shot)" model achieves a 0.48 score, outperforming the METEOR metric by 0.19, while our "Regressed GPT-4" model shows even greater alignment with expert evaluations, exceeding the best existing metric by a 0.35 margin. Moreover, the robustness of our explanations has been validated through a thorough iterative strategy. We plan to publicly release annotations from radiology experts, setting a new standard for accuracy in future assessments. This underscores the potential of our approach in enhancing the quality assessment of AI-driven medical reports.
format Preprint
id arxiv_https___arxiv_org_abs_2401_16578
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging Professional Radiologists' Expertise to Enhance LLMs' Evaluation for Radiology Reports
Zhu, Qingqing
Chen, Xiuying
Jin, Qiao
Hou, Benjamin
Mathai, Tejas Sudharshan
Mukherjee, Pritam
Gao, Xin
Summers, Ronald M
Lu, Zhiyong
Computation and Language
Artificial Intelligence
In radiology, Artificial Intelligence (AI) has significantly advanced report generation, but automatic evaluation of these AI-produced reports remains challenging. Current metrics, such as Conventional Natural Language Generation (NLG) and Clinical Efficacy (CE), often fall short in capturing the semantic intricacies of clinical contexts or overemphasize clinical details, undermining report clarity. To overcome these issues, our proposed method synergizes the expertise of professional radiologists with Large Language Models (LLMs), like GPT-3.5 and GPT-4 1. Utilizing In-Context Instruction Learning (ICIL) and Chain of Thought (CoT) reasoning, our approach aligns LLM evaluations with radiologist standards, enabling detailed comparisons between human and AI generated reports. This is further enhanced by a Regression model that aggregates sentence evaluation scores. Experimental results show that our "Detailed GPT-4 (5-shot)" model achieves a 0.48 score, outperforming the METEOR metric by 0.19, while our "Regressed GPT-4" model shows even greater alignment with expert evaluations, exceeding the best existing metric by a 0.35 margin. Moreover, the robustness of our explanations has been validated through a thorough iterative strategy. We plan to publicly release annotations from radiology experts, setting a new standard for accuracy in future assessments. This underscores the potential of our approach in enhancing the quality assessment of AI-driven medical reports.
title Leveraging Professional Radiologists' Expertise to Enhance LLMs' Evaluation for Radiology Reports
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.16578