Protein Representation Learning by Capturing Protein Sequence-Structure-Function Relationship

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ko, Eunji, Lee, Seul, Kim, Minseon, Kim, Dongki
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914791452311552
author Ko, Eunji
Lee, Seul
Kim, Minseon
Kim, Dongki
author_facet Ko, Eunji
Lee, Seul
Kim, Minseon
Kim, Dongki
contents The goal of protein representation learning is to extract knowledge from protein databases that can be applied to various protein-related downstream tasks. Although protein sequence, structure, and function are the three key modalities for a comprehensive understanding of proteins, existing methods for protein representation learning have utilized only one or two of these modalities due to the difficulty of capturing the asymmetric interrelationships between them. To account for this asymmetry, we introduce our novel asymmetric multi-modal masked autoencoder (AMMA). AMMA adopts (1) a unified multi-modal encoder to integrate all three modalities into a unified representation space and (2) asymmetric decoders to ensure that sequence latent features reflect structural and functional information. The experiments demonstrate that the proposed AMMA is highly effective in learning protein representations that exhibit well-aligned inter-modal relationships, which in turn makes it effective for various downstream protein-related tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_06663
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Protein Representation Learning by Capturing Protein Sequence-Structure-Function Relationship
Ko, Eunji
Lee, Seul
Kim, Minseon
Kim, Dongki
Biomolecules
Artificial Intelligence
Machine Learning
The goal of protein representation learning is to extract knowledge from protein databases that can be applied to various protein-related downstream tasks. Although protein sequence, structure, and function are the three key modalities for a comprehensive understanding of proteins, existing methods for protein representation learning have utilized only one or two of these modalities due to the difficulty of capturing the asymmetric interrelationships between them. To account for this asymmetry, we introduce our novel asymmetric multi-modal masked autoencoder (AMMA). AMMA adopts (1) a unified multi-modal encoder to integrate all three modalities into a unified representation space and (2) asymmetric decoders to ensure that sequence latent features reflect structural and functional information. The experiments demonstrate that the proposed AMMA is highly effective in learning protein representations that exhibit well-aligned inter-modal relationships, which in turn makes it effective for various downstream protein-related tasks.
title Protein Representation Learning by Capturing Protein Sequence-Structure-Function Relationship
topic Biomolecules
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2405.06663