Saved in:
Bibliographic Details
Main Authors: Xu, Tianhan, Hu, Zhe, Chen, Ling, Li, Bin
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2402.00474
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910314094657536
author Xu, Tianhan
Hu, Zhe
Chen, Ling
Li, Bin
author_facet Xu, Tianhan
Hu, Zhe
Chen, Ling
Li, Bin
contents Recent advances in large language models (LLMs) have demonstrated exceptional performance in various natural language processing (NLP) tasks. However, their effective application in the medical domain is hampered by a lack of medical domain knowledge. In this study, we present SA-MDKIF, a scalable and adaptable framework that aims to inject medical knowledge into general-purpose LLMs through instruction tuning, thereby enabling adaptability for various downstream tasks. SA-MDKIF consists of two stages: skill training and skill adaptation. In the first stage, we define 12 basic medical skills and use AdaLoRA to train these skills based on uniformly formatted instructional datasets that we have constructed. In the next stage, we train the skill router using task-specific downstream data and use this router to integrate the acquired skills with LLMs during inference. Experimental results on 9 different medical tasks show that SA-MDKIF improves performance by 10-20% compared to the original LLMs. Notably, this improvement is particularly pronounced for unseen medical tasks, showing an improvement of up to 30%.
format Preprint
id arxiv_https___arxiv_org_abs_2402_00474
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SA-MDKIF: A Scalable and Adaptable Medical Domain Knowledge Injection Framework for Large Language Models
Xu, Tianhan
Hu, Zhe
Chen, Ling
Li, Bin
Computation and Language
Artificial Intelligence
Recent advances in large language models (LLMs) have demonstrated exceptional performance in various natural language processing (NLP) tasks. However, their effective application in the medical domain is hampered by a lack of medical domain knowledge. In this study, we present SA-MDKIF, a scalable and adaptable framework that aims to inject medical knowledge into general-purpose LLMs through instruction tuning, thereby enabling adaptability for various downstream tasks. SA-MDKIF consists of two stages: skill training and skill adaptation. In the first stage, we define 12 basic medical skills and use AdaLoRA to train these skills based on uniformly formatted instructional datasets that we have constructed. In the next stage, we train the skill router using task-specific downstream data and use this router to integrate the acquired skills with LLMs during inference. Experimental results on 9 different medical tasks show that SA-MDKIF improves performance by 10-20% compared to the original LLMs. Notably, this improvement is particularly pronounced for unseen medical tasks, showing an improvement of up to 30%.
title SA-MDKIF: A Scalable and Adaptable Medical Domain Knowledge Injection Framework for Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2402.00474