Benchmarking large language models for biomedical natural language processing applications and recommendations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Chen, Qingyu, Hu, Yan, Peng, Xueqing, Xie, Qianqian, Jin, Qiao, Gilson, Aidan, Singer, Maxwell B., Ai, Xuguang, Lai, Po-Ting, Wang, Zhizheng, Keloth, Vipina Kuttichi, Raja, Kalpana, Huang, Jiming, He, Huan, Lin, Fongci, Du, Jingcheng, Zhang, Rui, Zheng, W. Jim, Adelman, Ron A., Lu, Zhiyong, Xu, Hua
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909594584875008
author Chen, Qingyu
Hu, Yan
Peng, Xueqing
Xie, Qianqian
Jin, Qiao
Gilson, Aidan
Singer, Maxwell B.
Ai, Xuguang
Lai, Po-Ting
Wang, Zhizheng
Keloth, Vipina Kuttichi
Raja, Kalpana
Huang, Jiming
He, Huan
Lin, Fongci
Du, Jingcheng
Zhang, Rui
Zheng, W. Jim
Adelman, Ron A.
Lu, Zhiyong
Xu, Hua
author_facet Chen, Qingyu
Hu, Yan
Peng, Xueqing
Xie, Qianqian
Jin, Qiao
Gilson, Aidan
Singer, Maxwell B.
Ai, Xuguang
Lai, Po-Ting
Wang, Zhizheng
Keloth, Vipina Kuttichi
Raja, Kalpana
Huang, Jiming
He, Huan
Lin, Fongci
Du, Jingcheng
Zhang, Rui
Zheng, W. Jim
Adelman, Ron A.
Lu, Zhiyong
Xu, Hua
contents The rapid growth of biomedical literature poses challenges for manual knowledge curation and synthesis. Biomedical Natural Language Processing (BioNLP) automates the process. While Large Language Models (LLMs) have shown promise in general domains, their effectiveness in BioNLP tasks remains unclear due to limited benchmarks and practical guidelines. We perform a systematic evaluation of four LLMs, GPT and LLaMA representatives on 12 BioNLP benchmarks across six applications. We compare their zero-shot, few-shot, and fine-tuning performance with traditional fine-tuning of BERT or BART models. We examine inconsistencies, missing information, hallucinations, and perform cost analysis. Here we show that traditional fine-tuning outperforms zero or few shot LLMs in most tasks. However, closed-source LLMs like GPT-4 excel in reasoning-related tasks such as medical question answering. Open source LLMs still require fine-tuning to close performance gaps. We find issues like missing information and hallucinations in LLM outputs. These results offer practical insights for applying LLMs in BioNLP.
format Preprint
id arxiv_https___arxiv_org_abs_2305_16326
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Benchmarking large language models for biomedical natural language processing applications and recommendations
Chen, Qingyu
Hu, Yan
Peng, Xueqing
Xie, Qianqian
Jin, Qiao
Gilson, Aidan
Singer, Maxwell B.
Ai, Xuguang
Lai, Po-Ting
Wang, Zhizheng
Keloth, Vipina Kuttichi
Raja, Kalpana
Huang, Jiming
He, Huan
Lin, Fongci
Du, Jingcheng
Zhang, Rui
Zheng, W. Jim
Adelman, Ron A.
Lu, Zhiyong
Xu, Hua
Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
The rapid growth of biomedical literature poses challenges for manual knowledge curation and synthesis. Biomedical Natural Language Processing (BioNLP) automates the process. While Large Language Models (LLMs) have shown promise in general domains, their effectiveness in BioNLP tasks remains unclear due to limited benchmarks and practical guidelines. We perform a systematic evaluation of four LLMs, GPT and LLaMA representatives on 12 BioNLP benchmarks across six applications. We compare their zero-shot, few-shot, and fine-tuning performance with traditional fine-tuning of BERT or BART models. We examine inconsistencies, missing information, hallucinations, and perform cost analysis. Here we show that traditional fine-tuning outperforms zero or few shot LLMs in most tasks. However, closed-source LLMs like GPT-4 excel in reasoning-related tasks such as medical question answering. Open source LLMs still require fine-tuning to close performance gaps. We find issues like missing information and hallucinations in LLM outputs. These results offer practical insights for applying LLMs in BioNLP.
title Benchmarking large language models for biomedical natural language processing applications and recommendations
topic Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2305.16326