Automating Exploratory Proteomics Research via Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ding, Ning, Qu, Shang, Xie, Linhai, Li, Yifei, Liu, Zaoqu, Zhang, Kaiyan, Xiong, Yibai, Zuo, Yuxin, Chen, Zhangren, Hua, Ermo, Lv, Xingtai, Sun, Youbang, Li, Yang, Li, Dong, He, Fuchu, Zhou, Bowen
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913572733321216
author Ding, Ning
Qu, Shang
Xie, Linhai
Li, Yifei
Liu, Zaoqu
Zhang, Kaiyan
Xiong, Yibai
Zuo, Yuxin
Chen, Zhangren
Hua, Ermo
Lv, Xingtai
Sun, Youbang
Li, Yang
Li, Dong
He, Fuchu
Zhou, Bowen
author_facet Ding, Ning
Qu, Shang
Xie, Linhai
Li, Yifei
Liu, Zaoqu
Zhang, Kaiyan
Xiong, Yibai
Zuo, Yuxin
Chen, Zhangren
Hua, Ermo
Lv, Xingtai
Sun, Youbang
Li, Yang
Li, Dong
He, Fuchu
Zhou, Bowen
contents With the development of artificial intelligence, its contribution to science is evolving from simulating a complex problem to automating entire research processes and producing novel discoveries. Achieving this advancement requires both specialized general models grounded in real-world scientific data and iterative, exploratory frameworks that mirror human scientific methodologies. In this paper, we present PROTEUS, a fully automated system for scientific discovery from raw proteomics data. PROTEUS uses large language models (LLMs) to perform hierarchical planning, execute specialized bioinformatics tools, and iteratively refine analysis workflows to generate high-quality scientific hypotheses. The system takes proteomics datasets as input and produces a comprehensive set of research objectives, analysis results, and novel biological hypotheses without human intervention. We evaluated PROTEUS on 12 proteomics datasets collected from various biological samples (e.g. immune cells, tumors) and different sample types (single-cell and bulk), generating 191 scientific hypotheses. These were assessed using both automatic LLM-based scoring on 5 metrics and detailed reviews from human experts. Results demonstrate that PROTEUS consistently produces reliable, logically coherent results that align well with existing literature while also proposing novel, evaluable hypotheses. The system's flexible architecture facilitates seamless integration of diverse analysis tools and adaptation to different proteomics data types. By automating complex proteomics analysis workflows and hypothesis generation, PROTEUS has the potential to considerably accelerate the pace of scientific discovery in proteomics research, enabling researchers to efficiently explore large-scale datasets and uncover biological insights.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03743
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Automating Exploratory Proteomics Research via Language Models
Ding, Ning
Qu, Shang
Xie, Linhai
Li, Yifei
Liu, Zaoqu
Zhang, Kaiyan
Xiong, Yibai
Zuo, Yuxin
Chen, Zhangren
Hua, Ermo
Lv, Xingtai
Sun, Youbang
Li, Yang
Li, Dong
He, Fuchu
Zhou, Bowen
Artificial Intelligence
Quantitative Methods
With the development of artificial intelligence, its contribution to science is evolving from simulating a complex problem to automating entire research processes and producing novel discoveries. Achieving this advancement requires both specialized general models grounded in real-world scientific data and iterative, exploratory frameworks that mirror human scientific methodologies. In this paper, we present PROTEUS, a fully automated system for scientific discovery from raw proteomics data. PROTEUS uses large language models (LLMs) to perform hierarchical planning, execute specialized bioinformatics tools, and iteratively refine analysis workflows to generate high-quality scientific hypotheses. The system takes proteomics datasets as input and produces a comprehensive set of research objectives, analysis results, and novel biological hypotheses without human intervention. We evaluated PROTEUS on 12 proteomics datasets collected from various biological samples (e.g. immune cells, tumors) and different sample types (single-cell and bulk), generating 191 scientific hypotheses. These were assessed using both automatic LLM-based scoring on 5 metrics and detailed reviews from human experts. Results demonstrate that PROTEUS consistently produces reliable, logically coherent results that align well with existing literature while also proposing novel, evaluable hypotheses. The system's flexible architecture facilitates seamless integration of diverse analysis tools and adaptation to different proteomics data types. By automating complex proteomics analysis workflows and hypothesis generation, PROTEUS has the potential to considerably accelerate the pace of scientific discovery in proteomics research, enabling researchers to efficiently explore large-scale datasets and uncover biological insights.
title Automating Exploratory Proteomics Research via Language Models
topic Artificial Intelligence
Quantitative Methods
url https://arxiv.org/abs/2411.03743