Uncertainty-Guided Model Selection for Tabular Foundation Models in Biomolecule Efficacy Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jie, McCarthy, Andrew, Zhang, Zhizhuo, Young, Stephen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918155399462912
author Li, Jie
McCarthy, Andrew
Zhang, Zhizhuo
Young, Stephen
author_facet Li, Jie
McCarthy, Andrew
Zhang, Zhizhuo
Young, Stephen
contents In-context learners like TabPFN are promising for biomolecule efficacy prediction, where established molecular feature sets and relevant experimental results can serve as powerful contextual examples. However, their performance is highly sensitive to the provided context, making strategies like post-hoc ensembling of models trained on different data subsets a viable approach. An open question is how to select the best models for the ensemble without access to ground truth labels. In this study, we investigate an uncertainty-guided strategy for model selection. We demonstrate on an siRNA knockdown efficacy task that a TabPFN model using straightforward sequence-based features can surpass specialized state-of-the-art predictors. We also show that the model's predicted inter-quantile range (IQR), a measure of its uncertainty, has a negative correlation with true prediction error. We developed the OligoICP method, which selects and averages an ensemble of models with the lowest mean IQR for siRNA efficacy prediction, achieving superior performance compared to naive ensembling or using a single model trained on all available data. This finding highlights model uncertainty as a powerful, label-free heuristic for optimizing biomolecule efficacy predictions.
format Preprint
id arxiv_https___arxiv_org_abs_2510_02476
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Uncertainty-Guided Model Selection for Tabular Foundation Models in Biomolecule Efficacy Prediction
Li, Jie
McCarthy, Andrew
Zhang, Zhizhuo
Young, Stephen
Machine Learning
Quantitative Methods
In-context learners like TabPFN are promising for biomolecule efficacy prediction, where established molecular feature sets and relevant experimental results can serve as powerful contextual examples. However, their performance is highly sensitive to the provided context, making strategies like post-hoc ensembling of models trained on different data subsets a viable approach. An open question is how to select the best models for the ensemble without access to ground truth labels. In this study, we investigate an uncertainty-guided strategy for model selection. We demonstrate on an siRNA knockdown efficacy task that a TabPFN model using straightforward sequence-based features can surpass specialized state-of-the-art predictors. We also show that the model's predicted inter-quantile range (IQR), a measure of its uncertainty, has a negative correlation with true prediction error. We developed the OligoICP method, which selects and averages an ensemble of models with the lowest mean IQR for siRNA efficacy prediction, achieving superior performance compared to naive ensembling or using a single model trained on all available data. This finding highlights model uncertainty as a powerful, label-free heuristic for optimizing biomolecule efficacy predictions.
title Uncertainty-Guided Model Selection for Tabular Foundation Models in Biomolecule Efficacy Prediction
topic Machine Learning
Quantitative Methods
url https://arxiv.org/abs/2510.02476