3M-Diffusion: Latent Multi-Modal Diffusion for Language-Guided Molecular Structure Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhu, Huaisheng, Xiao, Teng, Honavar, Vasant G
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929525210742784
author Zhu, Huaisheng
Xiao, Teng
Honavar, Vasant G
author_facet Zhu, Huaisheng
Xiao, Teng
Honavar, Vasant G
contents Generating molecular structures with desired properties is a critical task with broad applications in drug discovery and materials design. We propose 3M-Diffusion, a novel multi-modal molecular graph generation method, to generate diverse, ideally novel molecular structures with desired properties. 3M-Diffusion encodes molecular graphs into a graph latent space which it then aligns with the text space learned by encoder-based LLMs from textual descriptions. It then reconstructs the molecular structure and atomic attributes based on the given text descriptions using the molecule decoder. It then learns a probabilistic mapping from the text space to the latent molecular graph space using a diffusion model. The results of our extensive experiments on several datasets demonstrate that 3M-Diffusion can generate high-quality, novel and diverse molecular graphs that semantically match the textual description provided.
format Preprint
id arxiv_https___arxiv_org_abs_2403_07179
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle 3M-Diffusion: Latent Multi-Modal Diffusion for Language-Guided Molecular Structure Generation
Zhu, Huaisheng
Xiao, Teng
Honavar, Vasant G
Machine Learning
Computation and Language
Biomolecules
Generating molecular structures with desired properties is a critical task with broad applications in drug discovery and materials design. We propose 3M-Diffusion, a novel multi-modal molecular graph generation method, to generate diverse, ideally novel molecular structures with desired properties. 3M-Diffusion encodes molecular graphs into a graph latent space which it then aligns with the text space learned by encoder-based LLMs from textual descriptions. It then reconstructs the molecular structure and atomic attributes based on the given text descriptions using the molecule decoder. It then learns a probabilistic mapping from the text space to the latent molecular graph space using a diffusion model. The results of our extensive experiments on several datasets demonstrate that 3M-Diffusion can generate high-quality, novel and diverse molecular graphs that semantically match the textual description provided.
title 3M-Diffusion: Latent Multi-Modal Diffusion for Language-Guided Molecular Structure Generation
topic Machine Learning
Computation and Language
Biomolecules
url https://arxiv.org/abs/2403.07179