GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Liu, Pengfei, Ren, Yiming, Tao, Jun, Ren, Zhixiang
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910319291400192
author Liu, Pengfei
Ren, Yiming
Tao, Jun
Ren, Zhixiang
author_facet Liu, Pengfei
Ren, Yiming
Tao, Jun
Ren, Zhixiang
contents Large language models have made significant strides in natural language processing, enabling innovative applications in molecular science by processing textual representations of molecules. However, most existing language models cannot capture the rich information with complex molecular structures or images. In this paper, we introduce GIT-Mol, a multi-modal large language model that integrates the Graph, Image, and Text information. To facilitate the integration of multi-modal molecular data, we propose GIT-Former, a novel architecture that is capable of aligning all modalities into a unified latent space. We achieve a 5%-10% accuracy increase in properties prediction and a 20.2% boost in molecule generation validity compared to the baselines. With the any-to-language molecular translation strategy, our model has the potential to perform more downstream tasks, such as compound name recognition and chemical reaction prediction.
format Preprint
id arxiv_https___arxiv_org_abs_2308_06911
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text
Liu, Pengfei
Ren, Yiming
Tao, Jun
Ren, Zhixiang
Machine Learning
Computation and Language
Biomolecules
Large language models have made significant strides in natural language processing, enabling innovative applications in molecular science by processing textual representations of molecules. However, most existing language models cannot capture the rich information with complex molecular structures or images. In this paper, we introduce GIT-Mol, a multi-modal large language model that integrates the Graph, Image, and Text information. To facilitate the integration of multi-modal molecular data, we propose GIT-Former, a novel architecture that is capable of aligning all modalities into a unified latent space. We achieve a 5%-10% accuracy increase in properties prediction and a 20.2% boost in molecule generation validity compared to the baselines. With the any-to-language molecular translation strategy, our model has the potential to perform more downstream tasks, such as compound name recognition and chemical reaction prediction.
title GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text
topic Machine Learning
Computation and Language
Biomolecules
url https://arxiv.org/abs/2308.06911