Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Wei, Wang, Chao, Chen, Liyi, Yin, Mingze, Zhu, Yiheng, Fu, Kun, Ye, Jieping, Xiong, Hui, Wang, Zheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912401505386496
author Wu, Wei
Wang, Chao
Chen, Liyi
Yin, Mingze
Zhu, Yiheng
Fu, Kun
Ye, Jieping
Xiong, Hui
Wang, Zheng
author_facet Wu, Wei
Wang, Chao
Chen, Liyi
Yin, Mingze
Zhu, Yiheng
Fu, Kun
Ye, Jieping
Xiong, Hui
Wang, Zheng
contents Proteins, as essential biomolecules, play a central role in biological processes, including metabolic reactions and DNA replication. Accurate prediction of their properties and functions is crucial in biological applications. Recent development of protein language models (pLMs) with supervised fine tuning provides a promising solution to this problem. However, the fine-tuned model is tailored for particular downstream prediction task, and achieving general-purpose protein understanding remains a challenge. In this paper, we introduce Structure-Enhanced Protein Instruction Tuning (SEPIT) framework to bridge this gap. Our approach incorporates a novel structure-aware module into pLMs to enrich their structural knowledge, and subsequently integrates these enhanced pLMs with large language models (LLMs) to advance protein understanding. In this framework, we propose a novel instruction tuning pipeline. First, we warm up the enhanced pLMs using contrastive learning and structure denoising. Then, caption-based instructions are used to establish a basic understanding of proteins. Finally, we refine this understanding by employing a mixture of experts (MoEs) to capture more complex properties and functional information with the same number of activated parameters. Moreover, we construct the largest and most comprehensive protein instruction dataset to date, which allows us to train and evaluate the general-purpose protein understanding model. Extensive experiments on both open-ended generation and closed-set answer tasks demonstrate the superior performance of SEPIT over both closed-source general LLMs and open-source LLMs trained with protein knowledge.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03553
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs
Wu, Wei
Wang, Chao
Chen, Liyi
Yin, Mingze
Zhu, Yiheng
Fu, Kun
Ye, Jieping
Xiong, Hui
Wang, Zheng
Computation and Language
Biomolecules
Proteins, as essential biomolecules, play a central role in biological processes, including metabolic reactions and DNA replication. Accurate prediction of their properties and functions is crucial in biological applications. Recent development of protein language models (pLMs) with supervised fine tuning provides a promising solution to this problem. However, the fine-tuned model is tailored for particular downstream prediction task, and achieving general-purpose protein understanding remains a challenge. In this paper, we introduce Structure-Enhanced Protein Instruction Tuning (SEPIT) framework to bridge this gap. Our approach incorporates a novel structure-aware module into pLMs to enrich their structural knowledge, and subsequently integrates these enhanced pLMs with large language models (LLMs) to advance protein understanding. In this framework, we propose a novel instruction tuning pipeline. First, we warm up the enhanced pLMs using contrastive learning and structure denoising. Then, caption-based instructions are used to establish a basic understanding of proteins. Finally, we refine this understanding by employing a mixture of experts (MoEs) to capture more complex properties and functional information with the same number of activated parameters. Moreover, we construct the largest and most comprehensive protein instruction dataset to date, which allows us to train and evaluate the general-purpose protein understanding model. Extensive experiments on both open-ended generation and closed-set answer tasks demonstrate the superior performance of SEPIT over both closed-source general LLMs and open-source LLMs trained with protein knowledge.
title Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs
topic Computation and Language
Biomolecules
url https://arxiv.org/abs/2410.03553