ProtCLIP: Function-Informed Protein Multi-Modal Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Hanjing, Yin, Mingze, Wu, Wei, Li, Mingyang, Fu, Kun, Chen, Jintai, Wu, Jian, Wang, Zheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916544991199232
author Zhou, Hanjing
Yin, Mingze
Wu, Wei
Li, Mingyang
Fu, Kun
Chen, Jintai
Wu, Jian
Wang, Zheng
author_facet Zhou, Hanjing
Yin, Mingze
Wu, Wei
Li, Mingyang
Fu, Kun
Chen, Jintai
Wu, Jian
Wang, Zheng
contents Multi-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were still unable to replicate the extraordinary success of language-supervised visual foundation models due to the ineffective usage of aligned protein-text paired data and the lack of an effective function-informed pre-training paradigm. To address these issues, this paper curates a large-scale protein-text paired dataset called ProtAnno with a property-driven sampling strategy, and introduces a novel function-informed protein pre-training paradigm. Specifically, the sampling strategy determines selecting probability based on the sample confidence and property coverage, balancing the data quality and data quantity in face of large-scale noisy data. Furthermore, motivated by significance of the protein specific functional mechanism, the proposed paradigm explicitly model protein static and dynamic functional segments by two segment-wise pre-training objectives, injecting fine-grained information in a function-informed manner. Leveraging all these innovations, we develop ProtCLIP, a multi-modality foundation model that comprehensively represents function-aware protein embeddings. On 22 different protein benchmarks within 5 types, including protein functionality classification, mutation effect prediction, cross-modal transformation, semantic similarity inference and protein-protein interaction prediction, our ProtCLIP consistently achieves SOTA performance, with remarkable improvements of 75% on average in five cross-modal transformation benchmarks, 59.9% in GO-CC and 39.7% in GO-BP protein function prediction. The experimental results verify the extraordinary potential of ProtCLIP serving as the protein multi-modality foundation model.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20014
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ProtCLIP: Function-Informed Protein Multi-Modal Learning
Zhou, Hanjing
Yin, Mingze
Wu, Wei
Li, Mingyang
Fu, Kun
Chen, Jintai
Wu, Jian
Wang, Zheng
Machine Learning
Artificial Intelligence
Biomolecules
Multi-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were still unable to replicate the extraordinary success of language-supervised visual foundation models due to the ineffective usage of aligned protein-text paired data and the lack of an effective function-informed pre-training paradigm. To address these issues, this paper curates a large-scale protein-text paired dataset called ProtAnno with a property-driven sampling strategy, and introduces a novel function-informed protein pre-training paradigm. Specifically, the sampling strategy determines selecting probability based on the sample confidence and property coverage, balancing the data quality and data quantity in face of large-scale noisy data. Furthermore, motivated by significance of the protein specific functional mechanism, the proposed paradigm explicitly model protein static and dynamic functional segments by two segment-wise pre-training objectives, injecting fine-grained information in a function-informed manner. Leveraging all these innovations, we develop ProtCLIP, a multi-modality foundation model that comprehensively represents function-aware protein embeddings. On 22 different protein benchmarks within 5 types, including protein functionality classification, mutation effect prediction, cross-modal transformation, semantic similarity inference and protein-protein interaction prediction, our ProtCLIP consistently achieves SOTA performance, with remarkable improvements of 75% on average in five cross-modal transformation benchmarks, 59.9% in GO-CC and 39.7% in GO-BP protein function prediction. The experimental results verify the extraordinary potential of ProtCLIP serving as the protein multi-modality foundation model.
title ProtCLIP: Function-Informed Protein Multi-Modal Learning
topic Machine Learning
Artificial Intelligence
Biomolecules
url https://arxiv.org/abs/2412.20014