Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Acikgoz, Emre Can, İnce, Osman Batur, Bench, Rayene, Boz, Arda Anıl, Kesen, İlker, Erdem, Aykut, Erdem, Erkut
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914770501763072
author Acikgoz, Emre Can
İnce, Osman Batur
Bench, Rayene
Boz, Arda Anıl
Kesen, İlker
Erdem, Aykut
Erdem, Erkut
author_facet Acikgoz, Emre Can
İnce, Osman Batur
Bench, Rayene
Boz, Arda Anıl
Kesen, İlker
Erdem, Aykut
Erdem, Erkut
contents The integration of Large Language Models (LLMs) into healthcare promises to transform medical diagnostics, research, and patient care. Yet, the progression of medical LLMs faces obstacles such as complex training requirements, rigorous evaluation demands, and the dominance of proprietary models that restrict academic exploration. Transparent, comprehensive access to LLM resources is essential for advancing the field, fostering reproducibility, and encouraging innovation in healthcare AI. We present Hippocrates, an open-source LLM framework specifically developed for the medical domain. In stark contrast to previous efforts, it offers unrestricted access to its training datasets, codebase, checkpoints, and evaluation protocols. This open approach is designed to stimulate collaborative research, allowing the community to build upon, refine, and rigorously evaluate medical LLMs within a transparent ecosystem. Also, we introduce Hippo, a family of 7B models tailored for the medical domain, fine-tuned from Mistral and LLaMA2 through continual pre-training, instruction tuning, and reinforcement learning from human and AI feedback. Our models outperform existing open medical LLMs models by a large-margin, even surpassing models with 70B parameters. Through Hippocrates, we aspire to unlock the full potential of LLMs not just to advance medical knowledge and patient care but also to democratize the benefits of AI research in healthcare, making them available across the globe.
format Preprint
id arxiv_https___arxiv_org_abs_2404_16621
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare
Acikgoz, Emre Can
İnce, Osman Batur
Bench, Rayene
Boz, Arda Anıl
Kesen, İlker
Erdem, Aykut
Erdem, Erkut
Machine Learning
Artificial Intelligence
Computation and Language
The integration of Large Language Models (LLMs) into healthcare promises to transform medical diagnostics, research, and patient care. Yet, the progression of medical LLMs faces obstacles such as complex training requirements, rigorous evaluation demands, and the dominance of proprietary models that restrict academic exploration. Transparent, comprehensive access to LLM resources is essential for advancing the field, fostering reproducibility, and encouraging innovation in healthcare AI. We present Hippocrates, an open-source LLM framework specifically developed for the medical domain. In stark contrast to previous efforts, it offers unrestricted access to its training datasets, codebase, checkpoints, and evaluation protocols. This open approach is designed to stimulate collaborative research, allowing the community to build upon, refine, and rigorously evaluate medical LLMs within a transparent ecosystem. Also, we introduce Hippo, a family of 7B models tailored for the medical domain, fine-tuned from Mistral and LLaMA2 through continual pre-training, instruction tuning, and reinforcement learning from human and AI feedback. Our models outperform existing open medical LLMs models by a large-margin, even surpassing models with 70B parameters. Through Hippocrates, we aspire to unlock the full potential of LLMs not just to advance medical knowledge and patient care but also to democratize the benefits of AI research in healthcare, making them available across the globe.
title Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2404.16621