ProteinGPT: Multimodal LLM for Protein Property Prediction and Structure Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Yijia, Sun, Edward, Jin, Yiqiao, Wang, Qifan, Wang, Wei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908325105369088
author Xiao, Yijia
Sun, Edward
Jin, Yiqiao
Wang, Qifan
Wang, Wei
author_facet Xiao, Yijia
Sun, Edward
Jin, Yiqiao
Wang, Qifan
Wang, Wei
contents Understanding biological processes, drug development, and biotechnological advancements requires a detailed analysis of protein structures and functions, a task that is inherently complex and time-consuming in traditional protein research. To streamline this process, we introduce ProteinGPT, a state-of-the-art multimodal large language model for proteins that enables users to upload protein sequences and/or structures for comprehensive analysis and responsive inquiries. ProteinGPT integrates protein sequence and structure encoders with linear projection layers to ensure precise representation adaptation and leverages a large language model (LLM) to generate accurate, contextually relevant responses. To train ProteinGPT, we constructed a large-scale dataset of 132,092 proteins, each annotated with 20-30 property tags and 5-10 QA pairs per protein, and optimized the instruction-tuning process using GPT-4o. Experiments demonstrate that ProteinGPT effectively generates informative responses to protein-related questions, achieving high performance on both semantic and lexical metrics and significantly outperforming baseline models and general-purpose LLMs in understanding and responding to protein-related queries. Our code and data are available at https://github.com/ProteinGPT/ProteinGPT.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11363
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ProteinGPT: Multimodal LLM for Protein Property Prediction and Structure Understanding
Xiao, Yijia
Sun, Edward
Jin, Yiqiao
Wang, Qifan
Wang, Wei
Artificial Intelligence
Computational Engineering, Finance, and Science
Machine Learning
Biomolecules
Understanding biological processes, drug development, and biotechnological advancements requires a detailed analysis of protein structures and functions, a task that is inherently complex and time-consuming in traditional protein research. To streamline this process, we introduce ProteinGPT, a state-of-the-art multimodal large language model for proteins that enables users to upload protein sequences and/or structures for comprehensive analysis and responsive inquiries. ProteinGPT integrates protein sequence and structure encoders with linear projection layers to ensure precise representation adaptation and leverages a large language model (LLM) to generate accurate, contextually relevant responses. To train ProteinGPT, we constructed a large-scale dataset of 132,092 proteins, each annotated with 20-30 property tags and 5-10 QA pairs per protein, and optimized the instruction-tuning process using GPT-4o. Experiments demonstrate that ProteinGPT effectively generates informative responses to protein-related questions, achieving high performance on both semantic and lexical metrics and significantly outperforming baseline models and general-purpose LLMs in understanding and responding to protein-related queries. Our code and data are available at https://github.com/ProteinGPT/ProteinGPT.
title ProteinGPT: Multimodal LLM for Protein Property Prediction and Structure Understanding
topic Artificial Intelligence
Computational Engineering, Finance, and Science
Machine Learning
Biomolecules
url https://arxiv.org/abs/2408.11363