MedINST: Meta Dataset of Biomedical Instructions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Wenhan, Fang, Meng, Zhang, Zihan, Yin, Yu, Song, Zirui, Chen, Ling, Pechenizkiy, Mykola, Chen, Qingyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914976438943744
author Han, Wenhan
Fang, Meng
Zhang, Zihan
Yin, Yu
Song, Zirui
Chen, Ling
Pechenizkiy, Mykola
Chen, Qingyu
author_facet Han, Wenhan
Fang, Meng
Zhang, Zihan
Yin, Yu
Song, Zirui
Chen, Ling
Pechenizkiy, Mykola
Chen, Qingyu
contents The integration of large language model (LLM) techniques in the field of medical analysis has brought about significant advancements, yet the scarcity of large, diverse, and well-annotated datasets remains a major challenge. Medical data and tasks, which vary in format, size, and other parameters, require extensive preprocessing and standardization for effective use in training LLMs. To address these challenges, we introduce MedINST, the Meta Dataset of Biomedical Instructions, a novel multi-domain, multi-task instructional meta-dataset. MedINST comprises 133 biomedical NLP tasks and over 7 million training samples, making it the most comprehensive biomedical instruction dataset to date. Using MedINST as the meta dataset, we curate MedINST32, a challenging benchmark with different task difficulties aiming to evaluate LLMs' generalization ability. We fine-tune several LLMs on MedINST and evaluate on MedINST32, showcasing enhanced cross-task generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2410_13458
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MedINST: Meta Dataset of Biomedical Instructions
Han, Wenhan
Fang, Meng
Zhang, Zihan
Yin, Yu
Song, Zirui
Chen, Ling
Pechenizkiy, Mykola
Chen, Qingyu
Computation and Language
The integration of large language model (LLM) techniques in the field of medical analysis has brought about significant advancements, yet the scarcity of large, diverse, and well-annotated datasets remains a major challenge. Medical data and tasks, which vary in format, size, and other parameters, require extensive preprocessing and standardization for effective use in training LLMs. To address these challenges, we introduce MedINST, the Meta Dataset of Biomedical Instructions, a novel multi-domain, multi-task instructional meta-dataset. MedINST comprises 133 biomedical NLP tasks and over 7 million training samples, making it the most comprehensive biomedical instruction dataset to date. Using MedINST as the meta dataset, we curate MedINST32, a challenging benchmark with different task difficulties aiming to evaluate LLMs' generalization ability. We fine-tune several LLMs on MedINST and evaluate on MedINST32, showcasing enhanced cross-task generalization.
title MedINST: Meta Dataset of Biomedical Instructions
topic Computation and Language
url https://arxiv.org/abs/2410.13458