Evaluating and Enhancing Large Language Models Performance in Domain-specific Medicine: Osteoarthritis Management with DocOA

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xi, You, MingKe, Wang, Li, Liu, WeiZhi, Fu, Yu, Xu, Jie, Zhang, Shaoting, Chen, Gang, Li, Kang, Li, Jian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909081954942976
author Chen, Xi
You, MingKe
Wang, Li
Liu, WeiZhi
Fu, Yu
Xu, Jie
Zhang, Shaoting
Chen, Gang
Li, Kang
Li, Jian
author_facet Chen, Xi
You, MingKe
Wang, Li
Liu, WeiZhi
Fu, Yu
Xu, Jie
Zhang, Shaoting
Chen, Gang
Li, Kang
Li, Jian
contents The efficacy of large language models (LLMs) in domain-specific medicine, particularly for managing complex diseases such as osteoarthritis (OA), remains largely unexplored. This study focused on evaluating and enhancing the clinical capabilities of LLMs in specific domains, using osteoarthritis (OA) management as a case study. A domain specific benchmark framework was developed, which evaluate LLMs across a spectrum from domain-specific knowledge to clinical applications in real-world clinical scenarios. DocOA, a specialized LLM tailored for OA management that integrates retrieval-augmented generation (RAG) and instruction prompts, was developed. The study compared the performance of GPT-3.5, GPT-4, and a specialized assistant, DocOA, using objective and human evaluations. Results showed that general LLMs like GPT-3.5 and GPT-4 were less effective in the specialized domain of OA management, particularly in providing personalized treatment recommendations. However, DocOA showed significant improvements. This study introduces a novel benchmark framework which assesses the domain-specific abilities of LLMs in multiple aspects, highlights the limitations of generalized LLMs in clinical contexts, and demonstrates the potential of tailored approaches for developing domain-specific medical LLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2401_12998
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating and Enhancing Large Language Models Performance in Domain-specific Medicine: Osteoarthritis Management with DocOA
Chen, Xi
You, MingKe
Wang, Li
Liu, WeiZhi
Fu, Yu
Xu, Jie
Zhang, Shaoting
Chen, Gang
Li, Kang
Li, Jian
Computation and Language
Artificial Intelligence
The efficacy of large language models (LLMs) in domain-specific medicine, particularly for managing complex diseases such as osteoarthritis (OA), remains largely unexplored. This study focused on evaluating and enhancing the clinical capabilities of LLMs in specific domains, using osteoarthritis (OA) management as a case study. A domain specific benchmark framework was developed, which evaluate LLMs across a spectrum from domain-specific knowledge to clinical applications in real-world clinical scenarios. DocOA, a specialized LLM tailored for OA management that integrates retrieval-augmented generation (RAG) and instruction prompts, was developed. The study compared the performance of GPT-3.5, GPT-4, and a specialized assistant, DocOA, using objective and human evaluations. Results showed that general LLMs like GPT-3.5 and GPT-4 were less effective in the specialized domain of OA management, particularly in providing personalized treatment recommendations. However, DocOA showed significant improvements. This study introduces a novel benchmark framework which assesses the domain-specific abilities of LLMs in multiple aspects, highlights the limitations of generalized LLMs in clinical contexts, and demonstrates the potential of tailored approaches for developing domain-specific medical LLMs.
title Evaluating and Enhancing Large Language Models Performance in Domain-specific Medicine: Osteoarthritis Management with DocOA
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2401.12998