Evaluation of ChatGPT Family of Models for Biomedical Reasoning and Classification

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chen, Shan, Li, Yingya, Lu, Sheng, Van, Hoang, Aerts, Hugo JWL, Savova, Guergana K., Bitterman, Danielle S.
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916180818657280
author Chen, Shan
Li, Yingya
Lu, Sheng
Van, Hoang
Aerts, Hugo JWL
Savova, Guergana K.
Bitterman, Danielle S.
author_facet Chen, Shan
Li, Yingya
Lu, Sheng
Van, Hoang
Aerts, Hugo JWL
Savova, Guergana K.
Bitterman, Danielle S.
contents Recent advances in large language models (LLMs) have shown impressive ability in biomedical question-answering, but have not been adequately investigated for more specific biomedical applications. This study investigates the performance of LLMs such as the ChatGPT family of models (GPT-3.5s, GPT-4) in biomedical tasks beyond question-answering. Because no patient data can be passed to the OpenAI API public interface, we evaluated model performance with over 10000 samples as proxies for two fundamental tasks in the clinical domain - classification and reasoning. The first task is classifying whether statements of clinical and policy recommendations in scientific literature constitute health advice. The second task is causal relation detection from the biomedical literature. We compared LLMs with simpler models, such as bag-of-words (BoW) with logistic regression, and fine-tuned BioBERT models. Despite the excitement around viral ChatGPT, we found that fine-tuning for two fundamental NLP tasks remained the best strategy. The simple BoW model performed on par with the most complex LLM prompting. Prompt engineering required significant investment.
format Preprint
id arxiv_https___arxiv_org_abs_2304_02496
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Evaluation of ChatGPT Family of Models for Biomedical Reasoning and Classification
Chen, Shan
Li, Yingya
Lu, Sheng
Van, Hoang
Aerts, Hugo JWL
Savova, Guergana K.
Bitterman, Danielle S.
Computation and Language
Artificial Intelligence
Recent advances in large language models (LLMs) have shown impressive ability in biomedical question-answering, but have not been adequately investigated for more specific biomedical applications. This study investigates the performance of LLMs such as the ChatGPT family of models (GPT-3.5s, GPT-4) in biomedical tasks beyond question-answering. Because no patient data can be passed to the OpenAI API public interface, we evaluated model performance with over 10000 samples as proxies for two fundamental tasks in the clinical domain - classification and reasoning. The first task is classifying whether statements of clinical and policy recommendations in scientific literature constitute health advice. The second task is causal relation detection from the biomedical literature. We compared LLMs with simpler models, such as bag-of-words (BoW) with logistic regression, and fine-tuned BioBERT models. Despite the excitement around viral ChatGPT, we found that fine-tuning for two fundamental NLP tasks remained the best strategy. The simple BoW model performed on par with the most complex LLM prompting. Prompt engineering required significant investment.
title Evaluation of ChatGPT Family of Models for Biomedical Reasoning and Classification
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2304.02496