Repurposing Protein Language Models for Latent Flow-Based Fitness Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Arroyo, Amaru Caceres, Bogensperger, Lea, Allam, Ahmed, Krauthammer, Michael, Schindler, Konrad, Narnhofer, Dominik
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910009020907520
author Arroyo, Amaru Caceres
Bogensperger, Lea
Allam, Ahmed
Krauthammer, Michael
Schindler, Konrad
Narnhofer, Dominik
author_facet Arroyo, Amaru Caceres
Bogensperger, Lea
Allam, Ahmed
Krauthammer, Michael
Schindler, Konrad
Narnhofer, Dominik
contents Protein fitness optimization is challenged by a vast combinatorial landscape where high-fitness variants are extremely sparse. Many current methods either underperform or require computationally expensive gradient-based sampling. We present CHASE, a framework that repurposes the evolutionary knowledge of pretrained protein language models by compressing their embeddings into a compact latent space. By training a conditional flow-matching model with classifier-free guidance, we enable the direct generation of high-fitness variants without predictor-based guidance during the ODE sampling steps. CHASE achieves state-of-the-art performance on AAV and GFP protein design benchmarks. Finally, we show that bootstrapping with synthetic data can further enhance performance in data-constrained settings.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02425
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Repurposing Protein Language Models for Latent Flow-Based Fitness Optimization
Arroyo, Amaru Caceres
Bogensperger, Lea
Allam, Ahmed
Krauthammer, Michael
Schindler, Konrad
Narnhofer, Dominik
Machine Learning
Quantitative Methods
Protein fitness optimization is challenged by a vast combinatorial landscape where high-fitness variants are extremely sparse. Many current methods either underperform or require computationally expensive gradient-based sampling. We present CHASE, a framework that repurposes the evolutionary knowledge of pretrained protein language models by compressing their embeddings into a compact latent space. By training a conditional flow-matching model with classifier-free guidance, we enable the direct generation of high-fitness variants without predictor-based guidance during the ODE sampling steps. CHASE achieves state-of-the-art performance on AAV and GFP protein design benchmarks. Finally, we show that bootstrapping with synthetic data can further enhance performance in data-constrained settings.
title Repurposing Protein Language Models for Latent Flow-Based Fitness Optimization
topic Machine Learning
Quantitative Methods
url https://arxiv.org/abs/2602.02425