Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Vatsal, Shubham, Singh, Ayush
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912084806074368
author Vatsal, Shubham
Singh, Ayush
author_facet Vatsal, Shubham
Singh, Ayush
contents Large language models (LLMs) have shown remarkable performance on many tasks in different domains. However, their performance in closed-book biomedical machine reading comprehension (MRC) has not been evaluated in depth. In this work, we evaluate GPT on four closed-book biomedical MRC benchmarks. We experiment with different conventional prompting techniques as well as introduce our own novel prompting method. To solve some of the retrieval problems inherent to LLMs, we propose a prompting strategy named Implicit Retrieval Augmented Generation (RAG) that alleviates the need for using vector databases to retrieve important chunks in traditional RAG setups. Moreover, we report qualitative assessments on the natural language generation outputs from our approach. The results show that our new prompting technique is able to get the best performance in two out of four datasets and ranks second in rest of them. Experiments show that modern-day LLMs like GPT even in a zero-shot setting can outperform supervised models, leading to new state-of-the-art (SoTA) results on two of the benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18682
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
Vatsal, Shubham
Singh, Ayush
Computation and Language
Artificial Intelligence
Machine Learning
Large language models (LLMs) have shown remarkable performance on many tasks in different domains. However, their performance in closed-book biomedical machine reading comprehension (MRC) has not been evaluated in depth. In this work, we evaluate GPT on four closed-book biomedical MRC benchmarks. We experiment with different conventional prompting techniques as well as introduce our own novel prompting method. To solve some of the retrieval problems inherent to LLMs, we propose a prompting strategy named Implicit Retrieval Augmented Generation (RAG) that alleviates the need for using vector databases to retrieve important chunks in traditional RAG setups. Moreover, we report qualitative assessments on the natural language generation outputs from our approach. The results show that our new prompting technique is able to get the best performance in two out of four datasets and ranks second in rest of them. Experiments show that modern-day LLMs like GPT even in a zero-shot setting can outperform supervised models, leading to new state-of-the-art (SoTA) results on two of the benchmarks.
title Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2405.18682