3DS: Medical Domain Adaptation of LLMs via Decomposed Difficulty-based Data Selection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Hongxin, Fang, Yue, Zhu, Runchuan, Jiang, Xinke, Zhang, Jinyang, Xu, Yongxin, Chu, Xu, Zhao, Junfeng, Wang, Yasha
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912590850949120
author Ding, Hongxin
Fang, Yue
Zhu, Runchuan
Jiang, Xinke
Zhang, Jinyang
Xu, Yongxin
Chu, Xu
Zhao, Junfeng
Wang, Yasha
author_facet Ding, Hongxin
Fang, Yue
Zhu, Runchuan
Jiang, Xinke
Zhang, Jinyang
Xu, Yongxin
Chu, Xu
Zhao, Junfeng
Wang, Yasha
contents Large Language Models(LLMs) excel in general tasks but struggle in specialized domains like healthcare due to limited domain-specific knowledge.Supervised Fine-Tuning(SFT) data construction for domain adaptation often relies on heuristic methods, such as GPT-4 annotation or manual data selection, with a data-centric focus on presumed diverse, high-quality datasets. However, these methods overlook the model's inherent knowledge distribution, introducing noise, redundancy, and irrelevant data, leading to a mismatch between the selected data and the model's learning task, resulting in suboptimal performance. To address this, we propose a two-stage model-centric data selection framework, Decomposed Difficulty Data Selection (3DS), which aligns data with the model's knowledge distribution for optimized adaptation. In Stage1, we apply Prompt-Driven Data Selection via Explicit Alignment, where the the model filters irrelevant or redundant data based on its internal knowledge. In Stage2, we perform Decomposed Difficulty Data Selection, where data selection is guided by our defined difficulty decomposition, using three metrics: Instruction Understanding, Response Confidence, and Response Correctness. Additionally, an attention-based importance weighting mechanism captures token importance for more accurate difficulty calibration. This two-stage approach ensures the selected data is not only aligned with the model's knowledge and preferences but also appropriately challenging for the model to learn, leading to more effective and targeted domain adaptation. In the case study of the medical domain, our extensive experiments on real-world healthcare datasets demonstrate the superiority of 3DS over exisiting methods in accuracy by over 5.29%. Our dataset and code has been open-sourced at https://github.com/PuppyKnightUniversity/3DS.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10901
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 3DS: Medical Domain Adaptation of LLMs via Decomposed Difficulty-based Data Selection
Ding, Hongxin
Fang, Yue
Zhu, Runchuan
Jiang, Xinke
Zhang, Jinyang
Xu, Yongxin
Chu, Xu
Zhao, Junfeng
Wang, Yasha
Machine Learning
Artificial Intelligence
Computation and Language
Large Language Models(LLMs) excel in general tasks but struggle in specialized domains like healthcare due to limited domain-specific knowledge.Supervised Fine-Tuning(SFT) data construction for domain adaptation often relies on heuristic methods, such as GPT-4 annotation or manual data selection, with a data-centric focus on presumed diverse, high-quality datasets. However, these methods overlook the model's inherent knowledge distribution, introducing noise, redundancy, and irrelevant data, leading to a mismatch between the selected data and the model's learning task, resulting in suboptimal performance. To address this, we propose a two-stage model-centric data selection framework, Decomposed Difficulty Data Selection (3DS), which aligns data with the model's knowledge distribution for optimized adaptation. In Stage1, we apply Prompt-Driven Data Selection via Explicit Alignment, where the the model filters irrelevant or redundant data based on its internal knowledge. In Stage2, we perform Decomposed Difficulty Data Selection, where data selection is guided by our defined difficulty decomposition, using three metrics: Instruction Understanding, Response Confidence, and Response Correctness. Additionally, an attention-based importance weighting mechanism captures token importance for more accurate difficulty calibration. This two-stage approach ensures the selected data is not only aligned with the model's knowledge and preferences but also appropriately challenging for the model to learn, leading to more effective and targeted domain adaptation. In the case study of the medical domain, our extensive experiments on real-world healthcare datasets demonstrate the superiority of 3DS over exisiting methods in accuracy by over 5.29%. Our dataset and code has been open-sourced at https://github.com/PuppyKnightUniversity/3DS.
title 3DS: Medical Domain Adaptation of LLMs via Decomposed Difficulty-based Data Selection
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.10901