Multi-modal Molecule Structure-text Model for Text-based Retrieval and Editing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Shengchao, Nie, Weili, Wang, Chengpeng, Lu, Jiarui, Qiao, Zhuoran, Liu, Ling, Tang, Jian, Xiao, Chaowei, Anandkumar, Anima
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929227631165440
author Liu, Shengchao
Nie, Weili
Wang, Chengpeng
Lu, Jiarui
Qiao, Zhuoran
Liu, Ling
Tang, Jian
Xiao, Chaowei
Anandkumar, Anima
author_facet Liu, Shengchao
Nie, Weili
Wang, Chengpeng
Lu, Jiarui
Qiao, Zhuoran
Liu, Ling
Tang, Jian
Xiao, Chaowei
Anandkumar, Anima
contents There is increasing adoption of artificial intelligence in drug discovery. However, existing studies use machine learning to mainly utilize the chemical structures of molecules but ignore the vast textual knowledge available in chemistry. Incorporating textual knowledge enables us to realize new drug design objectives, adapt to text-based instructions and predict complex biological activities. Here we present a multi-modal molecule structure-text model, MoleculeSTM, by jointly learning molecules' chemical structures and textual descriptions via a contrastive learning strategy. To train MoleculeSTM, we construct a large multi-modal dataset, namely, PubChemSTM, with over 280,000 chemical structure-text pairs. To demonstrate the effectiveness and utility of MoleculeSTM, we design two challenging zero-shot tasks based on text instructions, including structure-text retrieval and molecule editing. MoleculeSTM has two main properties: open vocabulary and compositionality via natural language. In experiments, MoleculeSTM obtains the state-of-the-art generalization ability to novel biochemical concepts across various benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2212_10789
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Multi-modal Molecule Structure-text Model for Text-based Retrieval and Editing
Liu, Shengchao
Nie, Weili
Wang, Chengpeng
Lu, Jiarui
Qiao, Zhuoran
Liu, Ling
Tang, Jian
Xiao, Chaowei
Anandkumar, Anima
Machine Learning
Computation and Language
Quantitative Methods
There is increasing adoption of artificial intelligence in drug discovery. However, existing studies use machine learning to mainly utilize the chemical structures of molecules but ignore the vast textual knowledge available in chemistry. Incorporating textual knowledge enables us to realize new drug design objectives, adapt to text-based instructions and predict complex biological activities. Here we present a multi-modal molecule structure-text model, MoleculeSTM, by jointly learning molecules' chemical structures and textual descriptions via a contrastive learning strategy. To train MoleculeSTM, we construct a large multi-modal dataset, namely, PubChemSTM, with over 280,000 chemical structure-text pairs. To demonstrate the effectiveness and utility of MoleculeSTM, we design two challenging zero-shot tasks based on text instructions, including structure-text retrieval and molecule editing. MoleculeSTM has two main properties: open vocabulary and compositionality via natural language. In experiments, MoleculeSTM obtains the state-of-the-art generalization ability to novel biochemical concepts across various benchmarks.
title Multi-modal Molecule Structure-text Model for Text-based Retrieval and Editing
topic Machine Learning
Computation and Language
Quantitative Methods
url https://arxiv.org/abs/2212.10789