Generalization within in silico screening

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Loukas, Andreas, Kessel, Pan, Gligorijevic, Vladimir, Bonneau, Richard
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916332118736896
author Loukas, Andreas
Kessel, Pan
Gligorijevic, Vladimir
Bonneau, Richard
author_facet Loukas, Andreas
Kessel, Pan
Gligorijevic, Vladimir
Bonneau, Richard
contents In silico screening uses predictive models to select a batch of compounds with favorable properties from a library for experimental validation. Unlike conventional learning paradigms, success in this context is measured by the performance of the predictive model on the selected subset of compounds rather than the entire set of predictions. By extending learning theory, we show that the selectivity of the selection policy can significantly impact generalization, with a higher risk of errors occurring when exclusively selecting predicted positives and when targeting rare properties. Our analysis suggests a way to mitigate these challenges. We show that generalization can be markedly enhanced when considering a model's ability to predict the fraction of desired outcomes in a batch. This is promising, as the primary aim of screening is not necessarily to pinpoint the label of each compound individually, but rather to assemble a batch enriched for desirable compounds. Our theoretical insights are empirically validated across diverse tasks, architectures, and screening scenarios, underscoring their applicability.
format Preprint
id arxiv_https___arxiv_org_abs_2307_09379
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Generalization within in silico screening
Loukas, Andreas
Kessel, Pan
Gligorijevic, Vladimir
Bonneau, Richard
Machine Learning
In silico screening uses predictive models to select a batch of compounds with favorable properties from a library for experimental validation. Unlike conventional learning paradigms, success in this context is measured by the performance of the predictive model on the selected subset of compounds rather than the entire set of predictions. By extending learning theory, we show that the selectivity of the selection policy can significantly impact generalization, with a higher risk of errors occurring when exclusively selecting predicted positives and when targeting rare properties. Our analysis suggests a way to mitigate these challenges. We show that generalization can be markedly enhanced when considering a model's ability to predict the fraction of desired outcomes in a batch. This is promising, as the primary aim of screening is not necessarily to pinpoint the label of each compound individually, but rather to assemble a batch enriched for desirable compounds. Our theoretical insights are empirically validated across diverse tasks, architectures, and screening scenarios, underscoring their applicability.
title Generalization within in silico screening
topic Machine Learning
url https://arxiv.org/abs/2307.09379