ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type Neighborhoods

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kmicikiewicz, Michal, Fortuin, Vincent, Szczurek, Ewa
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911232835977216
author Kmicikiewicz, Michal
Fortuin, Vincent
Szczurek, Ewa
author_facet Kmicikiewicz, Michal
Fortuin, Vincent
Szczurek, Ewa
contents Designing protein sequences of both high fitness and novelty is a challenging task in data-efficient protein engineering. Exploration beyond wild-type neighborhoods often leads to biologically implausible sequences or relies on surrogate models that lose fidelity in novel regions. Here, we propose ProSpero, an active learning framework in which a frozen pre-trained generative model is guided by a surrogate updated from oracle feedback. By integrating fitness-relevant residue selection with biologically-constrained Sequential Monte Carlo sampling, our approach enables exploration beyond wild-type neighborhoods while preserving biological plausibility. We show that our framework remains effective even when the surrogate is misspecified. ProSpero consistently outperforms or matches existing methods across diverse protein engineering tasks, retrieving sequences of both high fitness and novelty.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22494
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type Neighborhoods
Kmicikiewicz, Michal
Fortuin, Vincent
Szczurek, Ewa
Machine Learning
Designing protein sequences of both high fitness and novelty is a challenging task in data-efficient protein engineering. Exploration beyond wild-type neighborhoods often leads to biologically implausible sequences or relies on surrogate models that lose fidelity in novel regions. Here, we propose ProSpero, an active learning framework in which a frozen pre-trained generative model is guided by a surrogate updated from oracle feedback. By integrating fitness-relevant residue selection with biologically-constrained Sequential Monte Carlo sampling, our approach enables exploration beyond wild-type neighborhoods while preserving biological plausibility. We show that our framework remains effective even when the surrogate is misspecified. ProSpero consistently outperforms or matches existing methods across diverse protein engineering tasks, retrieving sequences of both high fitness and novelty.
title ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type Neighborhoods
topic Machine Learning
url https://arxiv.org/abs/2505.22494