RadioRAG: Online Retrieval-augmented Generation for Radiology Question Answering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arasteh, Soroosh Tayebi, Lotfinia, Mahshad, Bressem, Keno, Siepmann, Robert, Adams, Lisa, Ferber, Dyke, Kuhl, Christiane, Kather, Jakob Nikolas, Nebelung, Sven, Truhn, Daniel
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909652064665600
author Arasteh, Soroosh Tayebi
Lotfinia, Mahshad
Bressem, Keno
Siepmann, Robert
Adams, Lisa
Ferber, Dyke
Kuhl, Christiane
Kather, Jakob Nikolas
Nebelung, Sven
Truhn, Daniel
author_facet Arasteh, Soroosh Tayebi
Lotfinia, Mahshad
Bressem, Keno
Siepmann, Robert
Adams, Lisa
Ferber, Dyke
Kuhl, Christiane
Kather, Jakob Nikolas
Nebelung, Sven
Truhn, Daniel
contents Large language models (LLMs) often generate outdated or inaccurate information based on static training datasets. Retrieval-augmented generation (RAG) mitigates this by integrating outside data sources. While previous RAG systems used pre-assembled, fixed databases with limited flexibility, we have developed Radiology RAG (RadioRAG), an end-to-end framework that retrieves data from authoritative radiologic online sources in real-time. We evaluate the diagnostic accuracy of various LLMs when answering radiology-specific questions with and without access to additional online information via RAG. Using 80 questions from the RSNA Case Collection across radiologic subspecialties and 24 additional expert-curated questions with reference standard answers, LLMs (GPT-3.5-turbo, GPT-4, Mistral-7B, Mixtral-8x7B, and Llama3 [8B and 70B]) were prompted with and without RadioRAG in a zero-shot inference scenario RadioRAG retrieved context-specific information from Radiopaedia in real-time. Accuracy was investigated. Statistical analyses were performed using bootstrapping. The results were further compared with human performance. RadioRAG improved diagnostic accuracy across most LLMs, with relative accuracy increases ranging up to 54% for different LLMs. It matched or exceeded non-RAG models and the human radiologist in question answering across radiologic subspecialties, particularly in breast imaging and emergency radiology. However, the degree of improvement varied among models; GPT-3.5-turbo and Mixtral-8x7B-instruct-v0.1 saw notable gains, while Mistral-7B-instruct-v0.2 showed no improvement, highlighting variability in RadioRAG's effectiveness. LLMs benefit when provided access to domain-specific data beyond their training data. RadioRAG shows potential to improve LLM accuracy and factuality in radiology question answering by integrating real-time domain-specific data.
format Preprint
id arxiv_https___arxiv_org_abs_2407_15621
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RadioRAG: Online Retrieval-augmented Generation for Radiology Question Answering
Arasteh, Soroosh Tayebi
Lotfinia, Mahshad
Bressem, Keno
Siepmann, Robert
Adams, Lisa
Ferber, Dyke
Kuhl, Christiane
Kather, Jakob Nikolas
Nebelung, Sven
Truhn, Daniel
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) often generate outdated or inaccurate information based on static training datasets. Retrieval-augmented generation (RAG) mitigates this by integrating outside data sources. While previous RAG systems used pre-assembled, fixed databases with limited flexibility, we have developed Radiology RAG (RadioRAG), an end-to-end framework that retrieves data from authoritative radiologic online sources in real-time. We evaluate the diagnostic accuracy of various LLMs when answering radiology-specific questions with and without access to additional online information via RAG. Using 80 questions from the RSNA Case Collection across radiologic subspecialties and 24 additional expert-curated questions with reference standard answers, LLMs (GPT-3.5-turbo, GPT-4, Mistral-7B, Mixtral-8x7B, and Llama3 [8B and 70B]) were prompted with and without RadioRAG in a zero-shot inference scenario RadioRAG retrieved context-specific information from Radiopaedia in real-time. Accuracy was investigated. Statistical analyses were performed using bootstrapping. The results were further compared with human performance. RadioRAG improved diagnostic accuracy across most LLMs, with relative accuracy increases ranging up to 54% for different LLMs. It matched or exceeded non-RAG models and the human radiologist in question answering across radiologic subspecialties, particularly in breast imaging and emergency radiology. However, the degree of improvement varied among models; GPT-3.5-turbo and Mixtral-8x7B-instruct-v0.1 saw notable gains, while Mistral-7B-instruct-v0.2 showed no improvement, highlighting variability in RadioRAG's effectiveness. LLMs benefit when provided access to domain-specific data beyond their training data. RadioRAG shows potential to improve LLM accuracy and factuality in radiology question answering by integrating real-time domain-specific data.
title RadioRAG: Online Retrieval-augmented Generation for Radiology Question Answering
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.15621