SPADE: Faster Drug Discovery by Learning from Sparse Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nandakumar, Rahul, Fauber, Ben, Chakrabarti, Deepayan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910195243810816
author Nandakumar, Rahul
Fauber, Ben
Chakrabarti, Deepayan
author_facet Nandakumar, Rahul
Fauber, Ben
Chakrabarti, Deepayan
contents Drug discovery seeks molecules (ligands) that bind strongly and selectively to a target protein. However, fewer than 5% of candidate ligands pass the bar for even the early stages of drug discovery. Furthermore, we want methods that work for novel proteins for which we have no prior data. Starting from scratch, we have to iteratively select and test candidate ligands such that we find enough ligands of the desired quality in as few tests as possible. Our proposed algorithm, named SPADE, introduces a novel approach to ligand selection that requires only 40 tests on average to find 10 high-quality ligands. In one-vs-one comparisons, SPADE outperforms deep learning and Bayesian optimization methods on more proteins, achieving median improvements of 7%-32% in sample efficiency. SPADE is also 10x faster than its closest competitor at scoring candidate drugs. Dataset and code is available at https://anonymous.4open.science/r/SPADE_Fast_Drug_Discovery_by_Learning_from_Sparse_Data-F028/README.md
format Preprint
id arxiv_https___arxiv_org_abs_2605_05370
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SPADE: Faster Drug Discovery by Learning from Sparse Data
Nandakumar, Rahul
Fauber, Ben
Chakrabarti, Deepayan
Machine Learning
Artificial Intelligence
Drug discovery seeks molecules (ligands) that bind strongly and selectively to a target protein. However, fewer than 5% of candidate ligands pass the bar for even the early stages of drug discovery. Furthermore, we want methods that work for novel proteins for which we have no prior data. Starting from scratch, we have to iteratively select and test candidate ligands such that we find enough ligands of the desired quality in as few tests as possible. Our proposed algorithm, named SPADE, introduces a novel approach to ligand selection that requires only 40 tests on average to find 10 high-quality ligands. In one-vs-one comparisons, SPADE outperforms deep learning and Bayesian optimization methods on more proteins, achieving median improvements of 7%-32% in sample efficiency. SPADE is also 10x faster than its closest competitor at scoring candidate drugs. Dataset and code is available at https://anonymous.4open.science/r/SPADE_Fast_Drug_Discovery_by_Learning_from_Sparse_Data-F028/README.md
title SPADE: Faster Drug Discovery by Learning from Sparse Data
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.05370