Aligning Proteins and Language: A Foundation Model for Protein Retrieval

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wu, Qifeng, Liu, Zhengzhe, Zhu, Han, Zhao, Yizhou, Kihara, Daisuke, Xu, Min
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913886869913600
author Wu, Qifeng
Liu, Zhengzhe
Zhu, Han
Zhao, Yizhou
Kihara, Daisuke
Xu, Min
author_facet Wu, Qifeng
Liu, Zhengzhe
Zhu, Han
Zhao, Yizhou
Kihara, Daisuke
Xu, Min
contents This paper aims to retrieve proteins with similar structures and semantics from large-scale protein dataset, facilitating the functional interpretation of protein structures derived by structural determination methods like cryo-Electron Microscopy (cryo-EM). Motivated by the recent progress of vision-language models (VLMs), we propose a CLIP-style framework for aligning 3D protein structures with functional annotations using contrastive learning. For model training, we propose a large-scale dataset of approximately 200,000 protein-caption pairs with rich functional descriptors. We evaluate our model in both in-domain and more challenging cross-database retrieval on Protein Data Bank (PDB) and Electron Microscopy Data Bank (EMDB) dataset, respectively. In both cases, our approach demonstrates promising zero-shot retrieval performance, highlighting the potential of multimodal foundation models for structure-function understanding in protein biology.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08023
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aligning Proteins and Language: A Foundation Model for Protein Retrieval
Wu, Qifeng
Liu, Zhengzhe
Zhu, Han
Zhao, Yizhou
Kihara, Daisuke
Xu, Min
Biomolecules
Artificial Intelligence
Computational Engineering, Finance, and Science
Computer Vision and Pattern Recognition
Machine Learning
This paper aims to retrieve proteins with similar structures and semantics from large-scale protein dataset, facilitating the functional interpretation of protein structures derived by structural determination methods like cryo-Electron Microscopy (cryo-EM). Motivated by the recent progress of vision-language models (VLMs), we propose a CLIP-style framework for aligning 3D protein structures with functional annotations using contrastive learning. For model training, we propose a large-scale dataset of approximately 200,000 protein-caption pairs with rich functional descriptors. We evaluate our model in both in-domain and more challenging cross-database retrieval on Protein Data Bank (PDB) and Electron Microscopy Data Bank (EMDB) dataset, respectively. In both cases, our approach demonstrates promising zero-shot retrieval performance, highlighting the potential of multimodal foundation models for structure-function understanding in protein biology.
title Aligning Proteins and Language: A Foundation Model for Protein Retrieval
topic Biomolecules
Artificial Intelligence
Computational Engineering, Finance, and Science
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2506.08023