Enhancing the efficiency of protein language models with minimal wet-lab data through few-shot learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Ziyi, Zhang, Liang, Yu, Yuanxi, Li, Mingchen, Hong, Liang, Tan, Pan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914664649064448
author Zhou, Ziyi
Zhang, Liang
Yu, Yuanxi
Li, Mingchen
Hong, Liang
Tan, Pan
author_facet Zhou, Ziyi
Zhang, Liang
Yu, Yuanxi
Li, Mingchen
Hong, Liang
Tan, Pan
contents Accurately modeling the protein fitness landscapes holds great importance for protein engineering. Recently, due to their capacity and representation ability, pre-trained protein language models have achieved state-of-the-art performance in predicting protein fitness without experimental data. However, their predictions are limited in accuracy as well as interpretability. Furthermore, such deep learning models require abundant labeled training examples for performance improvements, posing a practical barrier. In this work, we introduce FSFP, a training strategy that can effectively optimize protein language models under extreme data scarcity. By combining the techniques of meta-transfer learning, learning to rank, and parameter-efficient fine-tuning, FSFP can significantly boost the performance of various protein language models using merely tens of labeled single-site mutants from the target protein. The experiments across 87 deep mutational scanning datasets underscore its superiority over both unsupervised and supervised approaches, revealing its potential in facilitating AI-guided protein design.
format Preprint
id arxiv_https___arxiv_org_abs_2402_02004
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing the efficiency of protein language models with minimal wet-lab data through few-shot learning
Zhou, Ziyi
Zhang, Liang
Yu, Yuanxi
Li, Mingchen
Hong, Liang
Tan, Pan
Biomolecules
Accurately modeling the protein fitness landscapes holds great importance for protein engineering. Recently, due to their capacity and representation ability, pre-trained protein language models have achieved state-of-the-art performance in predicting protein fitness without experimental data. However, their predictions are limited in accuracy as well as interpretability. Furthermore, such deep learning models require abundant labeled training examples for performance improvements, posing a practical barrier. In this work, we introduce FSFP, a training strategy that can effectively optimize protein language models under extreme data scarcity. By combining the techniques of meta-transfer learning, learning to rank, and parameter-efficient fine-tuning, FSFP can significantly boost the performance of various protein language models using merely tens of labeled single-site mutants from the target protein. The experiments across 87 deep mutational scanning datasets underscore its superiority over both unsupervised and supervised approaches, revealing its potential in facilitating AI-guided protein design.
title Enhancing the efficiency of protein language models with minimal wet-lab data through few-shot learning
topic Biomolecules
url https://arxiv.org/abs/2402.02004