Learning Molecular Representation in a Cell

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Gang, Seal, Srijit, Arevalo, John, Liang, Zhenwen, Carpenter, Anne E., Jiang, Meng, Singh, Shantanu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910630175309824
author Liu, Gang
Seal, Srijit
Arevalo, John
Liang, Zhenwen
Carpenter, Anne E.
Jiang, Meng
Singh, Shantanu
author_facet Liu, Gang
Seal, Srijit
Arevalo, John
Liang, Zhenwen
Carpenter, Anne E.
Jiang, Meng
Singh, Shantanu
contents Predicting drug efficacy and safety in vivo requires information on biological responses (e.g., cell morphology and gene expression) to small molecule perturbations. However, current molecular representation learning methods do not provide a comprehensive view of cell states under these perturbations and struggle to remove noise, hindering model generalization. We introduce the Information Alignment (InfoAlign) approach to learn molecular representations through the information bottleneck method in cells. We integrate molecules and cellular response data as nodes into a context graph, connecting them with weighted edges based on chemical, biological, and computational criteria. For each molecule in a training batch, InfoAlign optimizes the encoder's latent representation with a minimality objective to discard redundant structural information. A sufficiency objective decodes the representation to align with different feature spaces from the molecule's neighborhood in the context graph. We demonstrate that the proposed sufficiency objective for alignment is tighter than existing encoder-based contrastive methods. Empirically, we validate representations from InfoAlign in two downstream applications: molecular property prediction against up to 27 baseline methods across four datasets, plus zero-shot molecule-morphology matching.
format Preprint
id arxiv_https___arxiv_org_abs_2406_12056
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning Molecular Representation in a Cell
Liu, Gang
Seal, Srijit
Arevalo, John
Liang, Zhenwen
Carpenter, Anne E.
Jiang, Meng
Singh, Shantanu
Machine Learning
Quantitative Methods
Predicting drug efficacy and safety in vivo requires information on biological responses (e.g., cell morphology and gene expression) to small molecule perturbations. However, current molecular representation learning methods do not provide a comprehensive view of cell states under these perturbations and struggle to remove noise, hindering model generalization. We introduce the Information Alignment (InfoAlign) approach to learn molecular representations through the information bottleneck method in cells. We integrate molecules and cellular response data as nodes into a context graph, connecting them with weighted edges based on chemical, biological, and computational criteria. For each molecule in a training batch, InfoAlign optimizes the encoder's latent representation with a minimality objective to discard redundant structural information. A sufficiency objective decodes the representation to align with different feature spaces from the molecule's neighborhood in the context graph. We demonstrate that the proposed sufficiency objective for alignment is tighter than existing encoder-based contrastive methods. Empirically, we validate representations from InfoAlign in two downstream applications: molecular property prediction against up to 27 baseline methods across four datasets, plus zero-shot molecule-morphology matching.
title Learning Molecular Representation in a Cell
topic Machine Learning
Quantitative Methods
url https://arxiv.org/abs/2406.12056