Advancing bioinformatics with large language models: components, applications and perspectives

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Liu, Jiajia, Yang, Mengyuan, Yu, Yankai, Xu, Haixia, Wang, Tiangang, Li, Kang, Zhou, Xiaobo
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913673455337472
author Liu, Jiajia
Yang, Mengyuan
Yu, Yankai
Xu, Haixia
Wang, Tiangang
Li, Kang
Zhou, Xiaobo
author_facet Liu, Jiajia
Yang, Mengyuan
Yu, Yankai
Xu, Haixia
Wang, Tiangang
Li, Kang
Zhou, Xiaobo
contents Large language models (LLMs) are a class of artificial intelligence models based on deep learning, which have great performance in various tasks, especially in natural language processing (NLP). Large language models typically consist of artificial neural networks with numerous parameters, trained on large amounts of unlabeled input using self-supervised or semi-supervised learning. However, their potential for solving bioinformatics problems may even exceed their proficiency in modeling human language. In this review, we will provide a comprehensive overview of the essential components of large language models (LLMs) in bioinformatics, spanning genomics, transcriptomics, proteomics, drug discovery, and single-cell analysis. Key aspects covered include tokenization methods for diverse data types, the architecture of transformer models, the core attention mechanism, and the pre-training processes underlying these models. Additionally, we will introduce currently available foundation models and highlight their downstream applications across various bioinformatics domains. Finally, drawing from our experience, we will offer practical guidance for both LLM users and developers, emphasizing strategies to optimize their use and foster further innovation in the field.
format Preprint
id arxiv_https___arxiv_org_abs_2401_04155
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Advancing bioinformatics with large language models: components, applications and perspectives
Liu, Jiajia
Yang, Mengyuan
Yu, Yankai
Xu, Haixia
Wang, Tiangang
Li, Kang
Zhou, Xiaobo
Quantitative Methods
Computation and Language
Large language models (LLMs) are a class of artificial intelligence models based on deep learning, which have great performance in various tasks, especially in natural language processing (NLP). Large language models typically consist of artificial neural networks with numerous parameters, trained on large amounts of unlabeled input using self-supervised or semi-supervised learning. However, their potential for solving bioinformatics problems may even exceed their proficiency in modeling human language. In this review, we will provide a comprehensive overview of the essential components of large language models (LLMs) in bioinformatics, spanning genomics, transcriptomics, proteomics, drug discovery, and single-cell analysis. Key aspects covered include tokenization methods for diverse data types, the architecture of transformer models, the core attention mechanism, and the pre-training processes underlying these models. Additionally, we will introduce currently available foundation models and highlight their downstream applications across various bioinformatics domains. Finally, drawing from our experience, we will offer practical guidance for both LLM users and developers, emphasizing strategies to optimize their use and foster further innovation in the field.
title Advancing bioinformatics with large language models: components, applications and perspectives
topic Quantitative Methods
Computation and Language
url https://arxiv.org/abs/2401.04155