Pretrained Joint Predictions for Scalable Batch Bayesian Optimization of Molecular Designs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang-Henderson, Miles, Kaufman, Benjamin, Williams, Edward, Pederson, Ryan, Rossi, Matteo, Howell, Owen, Underkoffler, Carl, Mardirossian, Narbe, Parkhill, John
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908652690997248
author Wang-Henderson, Miles
Kaufman, Benjamin
Williams, Edward
Pederson, Ryan
Rossi, Matteo
Howell, Owen
Underkoffler, Carl
Mardirossian, Narbe
Parkhill, John
author_facet Wang-Henderson, Miles
Kaufman, Benjamin
Williams, Edward
Pederson, Ryan
Rossi, Matteo
Howell, Owen
Underkoffler, Carl
Mardirossian, Narbe
Parkhill, John
contents Batched synthesis and testing of molecular designs is the key bottleneck of drug development. There has been great interest in leveraging biomolecular foundation models as surrogates to accelerate this process. In this work, we show how to obtain scalable probabilistic surrogates of binding affinity for use in Batch Bayesian Optimization (Batch BO). This demands parallel acquisition functions that hedge between designs and the ability to rapidly sample from a joint predictive density to approximate them. Through the framework of Epistemic Neural Networks (ENNs), we obtain scalable joint predictive distributions of binding affinity on top of representations taken from large structure-informed models. Key to this work is an investigation into the importance of prior networks in ENNs and how to pretrain them on synthetic data to improve downstream performance in Batch BO. Their utility is demonstrated by rediscovering known potent EGFR inhibitors on a semi-synthetic benchmark in up to 5x fewer iterations, as well as potent inhibitors from a real-world small-molecule library in up to 10x fewer iterations, offering a promising solution for large-scale drug discovery applications.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10590
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pretrained Joint Predictions for Scalable Batch Bayesian Optimization of Molecular Designs
Wang-Henderson, Miles
Kaufman, Benjamin
Williams, Edward
Pederson, Ryan
Rossi, Matteo
Howell, Owen
Underkoffler, Carl
Mardirossian, Narbe
Parkhill, John
Machine Learning
Batched synthesis and testing of molecular designs is the key bottleneck of drug development. There has been great interest in leveraging biomolecular foundation models as surrogates to accelerate this process. In this work, we show how to obtain scalable probabilistic surrogates of binding affinity for use in Batch Bayesian Optimization (Batch BO). This demands parallel acquisition functions that hedge between designs and the ability to rapidly sample from a joint predictive density to approximate them. Through the framework of Epistemic Neural Networks (ENNs), we obtain scalable joint predictive distributions of binding affinity on top of representations taken from large structure-informed models. Key to this work is an investigation into the importance of prior networks in ENNs and how to pretrain them on synthetic data to improve downstream performance in Batch BO. Their utility is demonstrated by rediscovering known potent EGFR inhibitors on a semi-synthetic benchmark in up to 5x fewer iterations, as well as potent inhibitors from a real-world small-molecule library in up to 10x fewer iterations, offering a promising solution for large-scale drug discovery applications.
title Pretrained Joint Predictions for Scalable Batch Bayesian Optimization of Molecular Designs
topic Machine Learning
url https://arxiv.org/abs/2511.10590