ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type Neighborhoods
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911232835977216 |
|---|---|
| author | Kmicikiewicz, Michal Fortuin, Vincent Szczurek, Ewa |
| author_facet | Kmicikiewicz, Michal Fortuin, Vincent Szczurek, Ewa |
| contents | Designing protein sequences of both high fitness and novelty is a challenging task in data-efficient protein engineering. Exploration beyond wild-type neighborhoods often leads to biologically implausible sequences or relies on surrogate models that lose fidelity in novel regions. Here, we propose ProSpero, an active learning framework in which a frozen pre-trained generative model is guided by a surrogate updated from oracle feedback. By integrating fitness-relevant residue selection with biologically-constrained Sequential Monte Carlo sampling, our approach enables exploration beyond wild-type neighborhoods while preserving biological plausibility. We show that our framework remains effective even when the surrogate is misspecified. ProSpero consistently outperforms or matches existing methods across diverse protein engineering tasks, retrieving sequences of both high fitness and novelty. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_22494 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type Neighborhoods Kmicikiewicz, Michal Fortuin, Vincent Szczurek, Ewa Machine Learning Designing protein sequences of both high fitness and novelty is a challenging task in data-efficient protein engineering. Exploration beyond wild-type neighborhoods often leads to biologically implausible sequences or relies on surrogate models that lose fidelity in novel regions. Here, we propose ProSpero, an active learning framework in which a frozen pre-trained generative model is guided by a surrogate updated from oracle feedback. By integrating fitness-relevant residue selection with biologically-constrained Sequential Monte Carlo sampling, our approach enables exploration beyond wild-type neighborhoods while preserving biological plausibility. We show that our framework remains effective even when the surrogate is misspecified. ProSpero consistently outperforms or matches existing methods across diverse protein engineering tasks, retrieving sequences of both high fitness and novelty. |
| title | ProSpero: Active Learning for Robust Protein Design Beyond Wild-Type Neighborhoods |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2505.22494 |