Y-Mol: A Multiscale Biomedical Knowledge-Guided Large Language Model for Drug Development

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Tengfei, Lin, Xuan, Li, Tianle, Li, Chaoyi, Chen, Long, Zhou, Peng, Cai, Xibao, Yang, Xinyu, Zeng, Daojian, Cao, Dongsheng, Zeng, Xiangxiang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916439918641152
author Ma, Tengfei
Lin, Xuan
Li, Tianle
Li, Chaoyi
Chen, Long
Zhou, Peng
Cai, Xibao
Yang, Xinyu
Zeng, Daojian
Cao, Dongsheng
Zeng, Xiangxiang
author_facet Ma, Tengfei
Lin, Xuan
Li, Tianle
Li, Chaoyi
Chen, Long
Zhou, Peng
Cai, Xibao
Yang, Xinyu
Zeng, Daojian
Cao, Dongsheng
Zeng, Xiangxiang
contents Large Language Models (LLMs) have recently demonstrated remarkable performance in general tasks across various fields. However, their effectiveness within specific domains such as drug development remains challenges. To solve these challenges, we introduce \textbf{Y-Mol}, forming a well-established LLM paradigm for the flow of drug development. Y-Mol is a multiscale biomedical knowledge-guided LLM designed to accomplish tasks across lead compound discovery, pre-clinic, and clinic prediction. By integrating millions of multiscale biomedical knowledge and using LLaMA2 as the base LLM, Y-Mol augments the reasoning capability in the biomedical domain by learning from a corpus of publications, knowledge graphs, and expert-designed synthetic data. The capability is further enriched with three types of drug-oriented instructions: description-based prompts from processed publications, semantic-based prompts for extracting associations from knowledge graphs, and template-based prompts for understanding expert knowledge from biomedical tools. Besides, Y-Mol offers a set of LLM paradigms that can autonomously execute the downstream tasks across the entire process of drug development, including virtual screening, drug design, pharmacological properties prediction, and drug-related interaction prediction. Our extensive evaluations of various biomedical sources demonstrate that Y-Mol significantly outperforms general-purpose LLMs in discovering lead compounds, predicting molecular properties, and identifying drug interaction events.
format Preprint
id arxiv_https___arxiv_org_abs_2410_11550
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Y-Mol: A Multiscale Biomedical Knowledge-Guided Large Language Model for Drug Development
Ma, Tengfei
Lin, Xuan
Li, Tianle
Li, Chaoyi
Chen, Long
Zhou, Peng
Cai, Xibao
Yang, Xinyu
Zeng, Daojian
Cao, Dongsheng
Zeng, Xiangxiang
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) have recently demonstrated remarkable performance in general tasks across various fields. However, their effectiveness within specific domains such as drug development remains challenges. To solve these challenges, we introduce \textbf{Y-Mol}, forming a well-established LLM paradigm for the flow of drug development. Y-Mol is a multiscale biomedical knowledge-guided LLM designed to accomplish tasks across lead compound discovery, pre-clinic, and clinic prediction. By integrating millions of multiscale biomedical knowledge and using LLaMA2 as the base LLM, Y-Mol augments the reasoning capability in the biomedical domain by learning from a corpus of publications, knowledge graphs, and expert-designed synthetic data. The capability is further enriched with three types of drug-oriented instructions: description-based prompts from processed publications, semantic-based prompts for extracting associations from knowledge graphs, and template-based prompts for understanding expert knowledge from biomedical tools. Besides, Y-Mol offers a set of LLM paradigms that can autonomously execute the downstream tasks across the entire process of drug development, including virtual screening, drug design, pharmacological properties prediction, and drug-related interaction prediction. Our extensive evaluations of various biomedical sources demonstrate that Y-Mol significantly outperforms general-purpose LLMs in discovering lead compounds, predicting molecular properties, and identifying drug interaction events.
title Y-Mol: A Multiscale Biomedical Knowledge-Guided Large Language Model for Drug Development
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2410.11550