Enhancing Uncertainty Quantification in Drug Discovery with Censored Regression Labels

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Svensson, Emma, Friesacher, Hannah Rosa, Winiwarter, Susanne, Mervin, Lewis, Arany, Adam, Engkvist, Ola
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917770348724224
author Svensson, Emma
Friesacher, Hannah Rosa
Winiwarter, Susanne
Mervin, Lewis
Arany, Adam
Engkvist, Ola
author_facet Svensson, Emma
Friesacher, Hannah Rosa
Winiwarter, Susanne
Mervin, Lewis
Arany, Adam
Engkvist, Ola
contents In the early stages of drug discovery, decisions regarding which experiments to pursue can be influenced by computational models. These decisions are critical due to the time-consuming and expensive nature of the experiments. Therefore, it is becoming essential to accurately quantify the uncertainty in machine learning predictions, such that resources can be used optimally and trust in the models improves. While computational methods for drug discovery often suffer from limited data and sparse experimental observations, additional information can exist in the form of censored labels that provide thresholds rather than precise values of observations. However, the standard approaches that quantify uncertainty in machine learning cannot fully utilize censored labels. In this work, we adapt ensemble-based, Bayesian, and Gaussian models with tools to learn from censored labels by using the Tobit model from survival analysis. Our results demonstrate that despite the partial information available in censored labels, they are essential to accurately and reliably model the real pharmaceutical setting.
format Preprint
id arxiv_https___arxiv_org_abs_2409_04313
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Uncertainty Quantification in Drug Discovery with Censored Regression Labels
Svensson, Emma
Friesacher, Hannah Rosa
Winiwarter, Susanne
Mervin, Lewis
Arany, Adam
Engkvist, Ola
Machine Learning
In the early stages of drug discovery, decisions regarding which experiments to pursue can be influenced by computational models. These decisions are critical due to the time-consuming and expensive nature of the experiments. Therefore, it is becoming essential to accurately quantify the uncertainty in machine learning predictions, such that resources can be used optimally and trust in the models improves. While computational methods for drug discovery often suffer from limited data and sparse experimental observations, additional information can exist in the form of censored labels that provide thresholds rather than precise values of observations. However, the standard approaches that quantify uncertainty in machine learning cannot fully utilize censored labels. In this work, we adapt ensemble-based, Bayesian, and Gaussian models with tools to learn from censored labels by using the Tobit model from survival analysis. Our results demonstrate that despite the partial information available in censored labels, they are essential to accurately and reliably model the real pharmaceutical setting.
title Enhancing Uncertainty Quantification in Drug Discovery with Censored Regression Labels
topic Machine Learning
url https://arxiv.org/abs/2409.04313