Protriever: End-to-End Differentiable Protein Homology Search for Fitness Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Weitzman, Ruben, Groth, Peter Mørch, Van Niekerk, Lood, Otani, Aoi, Gal, Yarin, Marks, Debora, Notin, Pascal
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908402208210944
author Weitzman, Ruben
Groth, Peter Mørch
Van Niekerk, Lood
Otani, Aoi
Gal, Yarin
Marks, Debora
Notin, Pascal
author_facet Weitzman, Ruben
Groth, Peter Mørch
Van Niekerk, Lood
Otani, Aoi
Gal, Yarin
Marks, Debora
Notin, Pascal
contents Retrieving homologous protein sequences is essential for a broad range of protein modeling tasks such as fitness prediction, protein design, structure modeling, and protein-protein interactions. Traditional workflows have relied on a two-step process: first retrieving homologs via Multiple Sequence Alignments (MSA), then training models on one or more of these alignments. However, MSA-based retrieval is computationally expensive, struggles with highly divergent sequences or complex insertions & deletions patterns, and operates independently of the downstream modeling objective. We introduce Protriever, an end-to-end differentiable framework that learns to retrieve relevant homologs while simultaneously training for the target task. When applied to protein fitness prediction, Protriever achieves state-of-the-art performance compared to sequence-based models that rely on MSA-based homolog retrieval, while being two orders of magnitude faster through efficient vector search. Protriever is both architecture- and task-agnostic, and can flexibly adapt to different retrieval strategies and protein databases at inference time -- offering a scalable alternative to alignment-centric approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08954
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Protriever: End-to-End Differentiable Protein Homology Search for Fitness Prediction
Weitzman, Ruben
Groth, Peter Mørch
Van Niekerk, Lood
Otani, Aoi
Gal, Yarin
Marks, Debora
Notin, Pascal
Quantitative Methods
Machine Learning
Biomolecules
Retrieving homologous protein sequences is essential for a broad range of protein modeling tasks such as fitness prediction, protein design, structure modeling, and protein-protein interactions. Traditional workflows have relied on a two-step process: first retrieving homologs via Multiple Sequence Alignments (MSA), then training models on one or more of these alignments. However, MSA-based retrieval is computationally expensive, struggles with highly divergent sequences or complex insertions & deletions patterns, and operates independently of the downstream modeling objective. We introduce Protriever, an end-to-end differentiable framework that learns to retrieve relevant homologs while simultaneously training for the target task. When applied to protein fitness prediction, Protriever achieves state-of-the-art performance compared to sequence-based models that rely on MSA-based homolog retrieval, while being two orders of magnitude faster through efficient vector search. Protriever is both architecture- and task-agnostic, and can flexibly adapt to different retrieval strategies and protein databases at inference time -- offering a scalable alternative to alignment-centric approaches.
title Protriever: End-to-End Differentiable Protein Homology Search for Fitness Prediction
topic Quantitative Methods
Machine Learning
Biomolecules
url https://arxiv.org/abs/2506.08954