Data augmentation enables label-specific generation of homologous protein sequences

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Rosset, Lorenzo, Weigt, Martin, Zamponi, Francesco
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918100019970048
author Rosset, Lorenzo
Weigt, Martin
Zamponi, Francesco
author_facet Rosset, Lorenzo
Weigt, Martin
Zamponi, Francesco
contents Accurately annotating and controlling protein function from sequence data remains a major challenge, particularly within homologous families where annotated sequences are scarce and structural variation is minimal. We present a two-stage approach for semi-supervised functional annotation and conditional sequence generation in protein families using representation learning. First, we demonstrate that protein language models, pretrained on large and diverse sequence datasets and possibly finetuned via contrastive learning, provide embeddings that robustly capture fine-grained functional specificities, even with limited labeled data. Second, we use the inferred annotations to train a generative probabilistic model, an annotation-aware Restricted Boltzmann Machine, capable of producing synthetic sequences with prescribed functional labels. Across several protein families, we show that this approach achieves highly accurate annotation quality and supports the generation of functionally coherent sequences. Our findings underscore the power of combining self-supervised learning with light supervision to overcome data scarcity in protein function prediction and design.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15651
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data augmentation enables label-specific generation of homologous protein sequences
Rosset, Lorenzo
Weigt, Martin
Zamponi, Francesco
Quantitative Methods
Disordered Systems and Neural Networks
Accurately annotating and controlling protein function from sequence data remains a major challenge, particularly within homologous families where annotated sequences are scarce and structural variation is minimal. We present a two-stage approach for semi-supervised functional annotation and conditional sequence generation in protein families using representation learning. First, we demonstrate that protein language models, pretrained on large and diverse sequence datasets and possibly finetuned via contrastive learning, provide embeddings that robustly capture fine-grained functional specificities, even with limited labeled data. Second, we use the inferred annotations to train a generative probabilistic model, an annotation-aware Restricted Boltzmann Machine, capable of producing synthetic sequences with prescribed functional labels. Across several protein families, we show that this approach achieves highly accurate annotation quality and supports the generation of functionally coherent sequences. Our findings underscore the power of combining self-supervised learning with light supervision to overcome data scarcity in protein function prediction and design.
title Data augmentation enables label-specific generation of homologous protein sequences
topic Quantitative Methods
Disordered Systems and Neural Networks
url https://arxiv.org/abs/2507.15651