Steering Generative Models with Experimental Data for Protein Fitness Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Jason, Chu, Wenda, Khalil, Daniel, Astudillo, Raul, Wittmann, Bruce J., Arnold, Frances H., Yue, Yisong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912661755658240
author Yang, Jason
Chu, Wenda
Khalil, Daniel
Astudillo, Raul
Wittmann, Bruce J.
Arnold, Frances H.
Yue, Yisong
author_facet Yang, Jason
Chu, Wenda
Khalil, Daniel
Astudillo, Raul
Wittmann, Bruce J.
Arnold, Frances H.
Yue, Yisong
contents Protein fitness optimization involves finding a protein sequence that maximizes desired quantitative properties in a combinatorially large design space of possible sequences. Recent advances in steering protein generative models (e.g., diffusion models and language models) with labeled data offer a promising approach. However, most previous studies have optimized surrogate rewards and/or utilized large amounts of labeled data for steering, making it unclear how well existing methods perform and compare to each other in real-world optimization campaigns where fitness is measured through low-throughput wet-lab assays. In this study, we explore fitness optimization using small amounts (hundreds) of labeled sequence-fitness pairs and comprehensively evaluate strategies such as classifier guidance and posterior sampling for guiding generation from different discrete diffusion models of protein sequences. We also demonstrate how guidance can be integrated into adaptive sequence selection akin to Thompson sampling in Bayesian optimization, showing that plug-and-play guidance strategies offer advantages over alternatives such as reinforcement learning with protein language models. Overall, we provide practical insights into how to effectively steer modern generative models for next-generation protein fitness optimization.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15093
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Steering Generative Models with Experimental Data for Protein Fitness Optimization
Yang, Jason
Chu, Wenda
Khalil, Daniel
Astudillo, Raul
Wittmann, Bruce J.
Arnold, Frances H.
Yue, Yisong
Biomolecules
Machine Learning
Protein fitness optimization involves finding a protein sequence that maximizes desired quantitative properties in a combinatorially large design space of possible sequences. Recent advances in steering protein generative models (e.g., diffusion models and language models) with labeled data offer a promising approach. However, most previous studies have optimized surrogate rewards and/or utilized large amounts of labeled data for steering, making it unclear how well existing methods perform and compare to each other in real-world optimization campaigns where fitness is measured through low-throughput wet-lab assays. In this study, we explore fitness optimization using small amounts (hundreds) of labeled sequence-fitness pairs and comprehensively evaluate strategies such as classifier guidance and posterior sampling for guiding generation from different discrete diffusion models of protein sequences. We also demonstrate how guidance can be integrated into adaptive sequence selection akin to Thompson sampling in Bayesian optimization, showing that plug-and-play guidance strategies offer advantages over alternatives such as reinforcement learning with protein language models. Overall, we provide practical insights into how to effectively steer modern generative models for next-generation protein fitness optimization.
title Steering Generative Models with Experimental Data for Protein Fitness Optimization
topic Biomolecules
Machine Learning
url https://arxiv.org/abs/2505.15093