Robust Optimization in Protein Fitness Landscapes Using Reinforcement Learning in Latent Space

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Lee, Minji, Vecchietti, Luiz Felipe, Jung, Hyunkyu, Ro, Hyun Joo, Cha, Meeyoung, Kim, Ho Min
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911892206780416
author Lee, Minji
Vecchietti, Luiz Felipe
Jung, Hyunkyu
Ro, Hyun Joo
Cha, Meeyoung
Kim, Ho Min
author_facet Lee, Minji
Vecchietti, Luiz Felipe
Jung, Hyunkyu
Ro, Hyun Joo
Cha, Meeyoung
Kim, Ho Min
contents Proteins are complex molecules responsible for different functions in nature. Enhancing the functionality of proteins and cellular fitness can significantly impact various industries. However, protein optimization using computational methods remains challenging, especially when starting from low-fitness sequences. We propose LatProtRL, an optimization method to efficiently traverse a latent space learned by an encoder-decoder leveraging a large protein language model. To escape local optima, our optimization is modeled as a Markov decision process using reinforcement learning acting directly in latent space. We evaluate our approach on two important fitness optimization tasks, demonstrating its ability to achieve comparable or superior fitness over baseline methods. Our findings and in vitro evaluation show that the generated sequences can reach high-fitness regions, suggesting a substantial potential of LatProtRL in lab-in-the-loop scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18986
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Robust Optimization in Protein Fitness Landscapes Using Reinforcement Learning in Latent Space
Lee, Minji
Vecchietti, Luiz Felipe
Jung, Hyunkyu
Ro, Hyun Joo
Cha, Meeyoung
Kim, Ho Min
Machine Learning
Biomolecules
Quantitative Methods
Proteins are complex molecules responsible for different functions in nature. Enhancing the functionality of proteins and cellular fitness can significantly impact various industries. However, protein optimization using computational methods remains challenging, especially when starting from low-fitness sequences. We propose LatProtRL, an optimization method to efficiently traverse a latent space learned by an encoder-decoder leveraging a large protein language model. To escape local optima, our optimization is modeled as a Markov decision process using reinforcement learning acting directly in latent space. We evaluate our approach on two important fitness optimization tasks, demonstrating its ability to achieve comparable or superior fitness over baseline methods. Our findings and in vitro evaluation show that the generated sequences can reach high-fitness regions, suggesting a substantial potential of LatProtRL in lab-in-the-loop scenarios.
title Robust Optimization in Protein Fitness Landscapes Using Reinforcement Learning in Latent Space
topic Machine Learning
Biomolecules
Quantitative Methods
url https://arxiv.org/abs/2405.18986