Repurposing Protein Language Models for Latent Flow-Based Fitness Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910009020907520 |
|---|---|
| author | Arroyo, Amaru Caceres Bogensperger, Lea Allam, Ahmed Krauthammer, Michael Schindler, Konrad Narnhofer, Dominik |
| author_facet | Arroyo, Amaru Caceres Bogensperger, Lea Allam, Ahmed Krauthammer, Michael Schindler, Konrad Narnhofer, Dominik |
| contents | Protein fitness optimization is challenged by a vast combinatorial landscape where high-fitness variants are extremely sparse. Many current methods either underperform or require computationally expensive gradient-based sampling. We present CHASE, a framework that repurposes the evolutionary knowledge of pretrained protein language models by compressing their embeddings into a compact latent space. By training a conditional flow-matching model with classifier-free guidance, we enable the direct generation of high-fitness variants without predictor-based guidance during the ODE sampling steps. CHASE achieves state-of-the-art performance on AAV and GFP protein design benchmarks. Finally, we show that bootstrapping with synthetic data can further enhance performance in data-constrained settings. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_02425 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Repurposing Protein Language Models for Latent Flow-Based Fitness Optimization Arroyo, Amaru Caceres Bogensperger, Lea Allam, Ahmed Krauthammer, Michael Schindler, Konrad Narnhofer, Dominik Machine Learning Quantitative Methods Protein fitness optimization is challenged by a vast combinatorial landscape where high-fitness variants are extremely sparse. Many current methods either underperform or require computationally expensive gradient-based sampling. We present CHASE, a framework that repurposes the evolutionary knowledge of pretrained protein language models by compressing their embeddings into a compact latent space. By training a conditional flow-matching model with classifier-free guidance, we enable the direct generation of high-fitness variants without predictor-based guidance during the ODE sampling steps. CHASE achieves state-of-the-art performance on AAV and GFP protein design benchmarks. Finally, we show that bootstrapping with synthetic data can further enhance performance in data-constrained settings. |
| title | Repurposing Protein Language Models for Latent Flow-Based Fitness Optimization |
| topic | Machine Learning Quantitative Methods |
| url | https://arxiv.org/abs/2602.02425 |