Multi-view biomedical foundation models for molecule-target and property prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Suryanarayanan, Parthasarathy, Qiu, Yunguang, Sethi, Shreyans, Mahajan, Diwakar, Li, Hongyang, Yang, Yuxin, Eyigoz, Elif, Saenz, Aldo Guzman, Platt, Daniel E., Rumbell, Timothy H., Ng, Kenney, Dey, Sanjoy, Burch, Myson, Kwon, Bum Chul, Meyer, Pablo, Cheng, Feixiong, Hu, Jianying, Morrone, Joseph A.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914353181097984
author Suryanarayanan, Parthasarathy
Qiu, Yunguang
Sethi, Shreyans
Mahajan, Diwakar
Li, Hongyang
Yang, Yuxin
Eyigoz, Elif
Saenz, Aldo Guzman
Platt, Daniel E.
Rumbell, Timothy H.
Ng, Kenney
Dey, Sanjoy
Burch, Myson
Kwon, Bum Chul
Meyer, Pablo
Cheng, Feixiong
Hu, Jianying
Morrone, Joseph A.
author_facet Suryanarayanan, Parthasarathy
Qiu, Yunguang
Sethi, Shreyans
Mahajan, Diwakar
Li, Hongyang
Yang, Yuxin
Eyigoz, Elif
Saenz, Aldo Guzman
Platt, Daniel E.
Rumbell, Timothy H.
Ng, Kenney
Dey, Sanjoy
Burch, Myson
Kwon, Bum Chul
Meyer, Pablo
Cheng, Feixiong
Hu, Jianying
Morrone, Joseph A.
contents Quality molecular representations are key to foundation model development in bio-medical research. Previous efforts have typically focused on a single representation or molecular view, which may have strengths or weaknesses on a given task. We develop Multi-view Molecular Embedding with Late Fusion (MMELON), an approach that integrates graph, image and text views in a foundation model setting and may be readily extended to additional representations. Single-view foundation models are each pre-trained on a dataset of up to 200M molecules. The multi-view model performs robustly, matching the performance of the highest-ranked single-view. It is validated on over 120 tasks, including molecular solubility, ADME properties, and activity against G Protein-Coupled receptors (GPCRs). We identify 33 GPCRs that are related to Alzheimer's disease and employ the multi-view model to select strong binders from a compound screen. Predictions are validated through structure-based modeling and identification of key binding motifs.
format Preprint
id arxiv_https___arxiv_org_abs_2410_19704
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-view biomedical foundation models for molecule-target and property prediction
Suryanarayanan, Parthasarathy
Qiu, Yunguang
Sethi, Shreyans
Mahajan, Diwakar
Li, Hongyang
Yang, Yuxin
Eyigoz, Elif
Saenz, Aldo Guzman
Platt, Daniel E.
Rumbell, Timothy H.
Ng, Kenney
Dey, Sanjoy
Burch, Myson
Kwon, Bum Chul
Meyer, Pablo
Cheng, Feixiong
Hu, Jianying
Morrone, Joseph A.
Biomolecules
Artificial Intelligence
Machine Learning
Quality molecular representations are key to foundation model development in bio-medical research. Previous efforts have typically focused on a single representation or molecular view, which may have strengths or weaknesses on a given task. We develop Multi-view Molecular Embedding with Late Fusion (MMELON), an approach that integrates graph, image and text views in a foundation model setting and may be readily extended to additional representations. Single-view foundation models are each pre-trained on a dataset of up to 200M molecules. The multi-view model performs robustly, matching the performance of the highest-ranked single-view. It is validated on over 120 tasks, including molecular solubility, ADME properties, and activity against G Protein-Coupled receptors (GPCRs). We identify 33 GPCRs that are related to Alzheimer's disease and employ the multi-view model to select strong binders from a compound screen. Predictions are validated through structure-based modeling and identification of key binding motifs.
title Multi-view biomedical foundation models for molecule-target and property prediction
topic Biomolecules
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2410.19704