MutaPLM: Protein Language Modeling for Mutation Explanation and Engineering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Yizhen, Nie, Zikun, Hong, Massimo, Zhao, Suyuan, Zhou, Hao, Nie, Zaiqing
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916461325320192
author Luo, Yizhen
Nie, Zikun
Hong, Massimo
Zhao, Suyuan
Zhou, Hao
Nie, Zaiqing
author_facet Luo, Yizhen
Nie, Zikun
Hong, Massimo
Zhao, Suyuan
Zhou, Hao
Nie, Zaiqing
contents Studying protein mutations within amino acid sequences holds tremendous significance in life sciences. Protein language models (PLMs) have demonstrated strong capabilities in broad biological applications. However, due to architectural design and lack of supervision, PLMs model mutations implicitly with evolutionary plausibility, which is not satisfactory to serve as explainable and engineerable tools in real-world studies. To address these issues, we present MutaPLM, a unified framework for interpreting and navigating protein mutations with protein language models. MutaPLM introduces a protein delta network that captures explicit protein mutation representations within a unified feature space, and a transfer learning pipeline with a chain-of-thought (CoT) strategy to harvest protein mutation knowledge from biomedical texts. We also construct MutaDescribe, the first large-scale protein mutation dataset with rich textual annotations, which provides cross-modal supervision signals. Through comprehensive experiments, we demonstrate that MutaPLM excels at providing human-understandable explanations for mutational effects and prioritizing novel mutations with desirable properties. Our code, model, and data are open-sourced at https://github.com/PharMolix/MutaPLM.
format Preprint
id arxiv_https___arxiv_org_abs_2410_22949
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MutaPLM: Protein Language Modeling for Mutation Explanation and Engineering
Luo, Yizhen
Nie, Zikun
Hong, Massimo
Zhao, Suyuan
Zhou, Hao
Nie, Zaiqing
Machine Learning
Biomolecules
68T07
Studying protein mutations within amino acid sequences holds tremendous significance in life sciences. Protein language models (PLMs) have demonstrated strong capabilities in broad biological applications. However, due to architectural design and lack of supervision, PLMs model mutations implicitly with evolutionary plausibility, which is not satisfactory to serve as explainable and engineerable tools in real-world studies. To address these issues, we present MutaPLM, a unified framework for interpreting and navigating protein mutations with protein language models. MutaPLM introduces a protein delta network that captures explicit protein mutation representations within a unified feature space, and a transfer learning pipeline with a chain-of-thought (CoT) strategy to harvest protein mutation knowledge from biomedical texts. We also construct MutaDescribe, the first large-scale protein mutation dataset with rich textual annotations, which provides cross-modal supervision signals. Through comprehensive experiments, we demonstrate that MutaPLM excels at providing human-understandable explanations for mutational effects and prioritizing novel mutations with desirable properties. Our code, model, and data are open-sourced at https://github.com/PharMolix/MutaPLM.
title MutaPLM: Protein Language Modeling for Mutation Explanation and Engineering
topic Machine Learning
Biomolecules
68T07
url https://arxiv.org/abs/2410.22949