Contextual Representation Anchor Network to Alleviate Selection Bias in Few-Shot Drug Discovery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Ruifeng, Liu, Wei, Zhou, Xiangxin, Li, Mingqian, Zhang, Qiang, Chen, Hongyang, Lin, Xuemin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913566933647360
author Li, Ruifeng
Liu, Wei
Zhou, Xiangxin
Li, Mingqian
Zhang, Qiang
Chen, Hongyang
Lin, Xuemin
author_facet Li, Ruifeng
Liu, Wei
Zhou, Xiangxin
Li, Mingqian
Zhang, Qiang
Chen, Hongyang
Lin, Xuemin
contents In the drug discovery process, the low success rate of drug candidate screening often leads to insufficient labeled data, causing the few-shot learning problem in molecular property prediction. Existing methods for few-shot molecular property prediction overlook the sample selection bias, which arises from non-random sample selection in chemical experiments. This bias in data representativeness leads to suboptimal performance. To overcome this challenge, we present a novel method named contextual representation anchor Network (CRA), where an anchor refers to a cluster center of the representations of molecules and serves as a bridge to transfer enriched contextual knowledge into molecular representations and enhance their expressiveness. CRA introduces a dual-augmentation mechanism that includes context augmentation, which dynamically retrieves analogous unlabeled molecules and captures their task-specific contextual knowledge to enhance the anchors, and anchor augmentation, which leverages the anchors to augment the molecular representations. We evaluate our approach on the MoleculeNet and FS-Mol benchmarks, as well as in domain transfer experiments. The results demonstrate that CRA outperforms the state-of-the-art by 2.60% and 3.28% in AUC and $Δ$AUC-PR metrics, respectively, and exhibits superior generalization capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20711
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Contextual Representation Anchor Network to Alleviate Selection Bias in Few-Shot Drug Discovery
Li, Ruifeng
Liu, Wei
Zhou, Xiangxin
Li, Mingqian
Zhang, Qiang
Chen, Hongyang
Lin, Xuemin
Machine Learning
Artificial Intelligence
Biomolecules
68U07
I.2.1
In the drug discovery process, the low success rate of drug candidate screening often leads to insufficient labeled data, causing the few-shot learning problem in molecular property prediction. Existing methods for few-shot molecular property prediction overlook the sample selection bias, which arises from non-random sample selection in chemical experiments. This bias in data representativeness leads to suboptimal performance. To overcome this challenge, we present a novel method named contextual representation anchor Network (CRA), where an anchor refers to a cluster center of the representations of molecules and serves as a bridge to transfer enriched contextual knowledge into molecular representations and enhance their expressiveness. CRA introduces a dual-augmentation mechanism that includes context augmentation, which dynamically retrieves analogous unlabeled molecules and captures their task-specific contextual knowledge to enhance the anchors, and anchor augmentation, which leverages the anchors to augment the molecular representations. We evaluate our approach on the MoleculeNet and FS-Mol benchmarks, as well as in domain transfer experiments. The results demonstrate that CRA outperforms the state-of-the-art by 2.60% and 3.28% in AUC and $Δ$AUC-PR metrics, respectively, and exhibits superior generalization capabilities.
title Contextual Representation Anchor Network to Alleviate Selection Bias in Few-Shot Drug Discovery
topic Machine Learning
Artificial Intelligence
Biomolecules
68U07
I.2.1
url https://arxiv.org/abs/2410.20711