Robust Molecular Property Prediction via Densifying Scarce Labeled Data

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Jina, Willette, Jeffrey, Andreis, Bruno, Hwang, Sung Ju
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908742537183232
author Kim, Jina
Willette, Jeffrey
Andreis, Bruno
Hwang, Sung Ju
author_facet Kim, Jina
Willette, Jeffrey
Andreis, Bruno
Hwang, Sung Ju
contents A widely recognized limitation of molecular prediction models is their reliance on structures observed in the training data, resulting in poor generalization to out-of-distribution compounds. Yet in drug discovery, the compounds most critical for advancing research often lie beyond the training set, making the bias toward the training data particularly problematic. This mismatch introduces substantial covariate shift, under which standard deep learning models produce unstable and inaccurate predictions. Furthermore, the scarcity of labeled data-stemming from the onerous and costly nature of experimental validation-further exacerbates the difficulty of achieving reliable generalization. To address these limitations, we propose a novel bilevel optimization approach that leverages unlabeled data to interpolate between in-distribution (ID) and out-of-distribution (OOD) data, enabling the model to learn how to generalize beyond the training distribution. We demonstrate significant performance gains on challenging real-world datasets with substantial covariate shift, supported by t-SNE visualizations highlighting our interpolation method.
format Preprint
id arxiv_https___arxiv_org_abs_2506_11877
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robust Molecular Property Prediction via Densifying Scarce Labeled Data
Kim, Jina
Willette, Jeffrey
Andreis, Bruno
Hwang, Sung Ju
Machine Learning
Artificial Intelligence
A widely recognized limitation of molecular prediction models is their reliance on structures observed in the training data, resulting in poor generalization to out-of-distribution compounds. Yet in drug discovery, the compounds most critical for advancing research often lie beyond the training set, making the bias toward the training data particularly problematic. This mismatch introduces substantial covariate shift, under which standard deep learning models produce unstable and inaccurate predictions. Furthermore, the scarcity of labeled data-stemming from the onerous and costly nature of experimental validation-further exacerbates the difficulty of achieving reliable generalization. To address these limitations, we propose a novel bilevel optimization approach that leverages unlabeled data to interpolate between in-distribution (ID) and out-of-distribution (OOD) data, enabling the model to learn how to generalize beyond the training distribution. We demonstrate significant performance gains on challenging real-world datasets with substantial covariate shift, supported by t-SNE visualizations highlighting our interpolation method.
title Robust Molecular Property Prediction via Densifying Scarce Labeled Data
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.11877