CellVerse: Do Large Language Models Really Understand Cell Biology?

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Fan, Liu, Tianyu, Zhu, Zhihong, Wu, Hao, Wang, Haixin, Zhou, Donghao, Zheng, Yefeng, Wang, Kun, Wu, Xian, Heng, Pheng-Ann
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915285891547136
author Zhang, Fan
Liu, Tianyu
Zhu, Zhihong
Wu, Hao
Wang, Haixin
Zhou, Donghao
Zheng, Yefeng
Wang, Kun
Wu, Xian
Heng, Pheng-Ann
author_facet Zhang, Fan
Liu, Tianyu
Zhu, Zhihong
Wu, Hao
Wang, Haixin
Zhou, Donghao
Zheng, Yefeng
Wang, Kun
Wu, Xian
Heng, Pheng-Ann
contents Recent studies have demonstrated the feasibility of modeling single-cell data as natural languages and the potential of leveraging powerful large language models (LLMs) for understanding cell biology. However, a comprehensive evaluation of LLMs' performance on language-driven single-cell analysis tasks still remains unexplored. Motivated by this challenge, we introduce CellVerse, a unified language-centric question-answering benchmark that integrates four types of single-cell multi-omics data and encompasses three hierarchical levels of single-cell analysis tasks: cell type annotation (cell-level), drug response prediction (drug-level), and perturbation analysis (gene-level). Going beyond this, we systematically evaluate the performance across 14 open-source and closed-source LLMs ranging from 160M to 671B on CellVerse. Remarkably, the experimental results reveal: (1) Existing specialist models (C2S-Pythia) fail to make reasonable decisions across all sub-tasks within CellVerse, while generalist models such as Qwen, Llama, GPT, and DeepSeek family models exhibit preliminary understanding capabilities within the realm of cell biology. (2) The performance of current LLMs falls short of expectations and has substantial room for improvement. Notably, in the widely studied drug response prediction task, none of the evaluated LLMs demonstrate significant performance improvement over random guessing. CellVerse offers the first large-scale empirical demonstration that significant challenges still remain in applying LLMs to cell biology. By introducing CellVerse, we lay the foundation for advancing cell biology through natural languages and hope this paradigm could facilitate next-generation single-cell analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2505_07865
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CellVerse: Do Large Language Models Really Understand Cell Biology?
Zhang, Fan
Liu, Tianyu
Zhu, Zhihong
Wu, Hao
Wang, Haixin
Zhou, Donghao
Zheng, Yefeng
Wang, Kun
Wu, Xian
Heng, Pheng-Ann
Quantitative Methods
Artificial Intelligence
Computation and Language
Cell Behavior
Recent studies have demonstrated the feasibility of modeling single-cell data as natural languages and the potential of leveraging powerful large language models (LLMs) for understanding cell biology. However, a comprehensive evaluation of LLMs' performance on language-driven single-cell analysis tasks still remains unexplored. Motivated by this challenge, we introduce CellVerse, a unified language-centric question-answering benchmark that integrates four types of single-cell multi-omics data and encompasses three hierarchical levels of single-cell analysis tasks: cell type annotation (cell-level), drug response prediction (drug-level), and perturbation analysis (gene-level). Going beyond this, we systematically evaluate the performance across 14 open-source and closed-source LLMs ranging from 160M to 671B on CellVerse. Remarkably, the experimental results reveal: (1) Existing specialist models (C2S-Pythia) fail to make reasonable decisions across all sub-tasks within CellVerse, while generalist models such as Qwen, Llama, GPT, and DeepSeek family models exhibit preliminary understanding capabilities within the realm of cell biology. (2) The performance of current LLMs falls short of expectations and has substantial room for improvement. Notably, in the widely studied drug response prediction task, none of the evaluated LLMs demonstrate significant performance improvement over random guessing. CellVerse offers the first large-scale empirical demonstration that significant challenges still remain in applying LLMs to cell biology. By introducing CellVerse, we lay the foundation for advancing cell biology through natural languages and hope this paradigm could facilitate next-generation single-cell analysis.
title CellVerse: Do Large Language Models Really Understand Cell Biology?
topic Quantitative Methods
Artificial Intelligence
Computation and Language
Cell Behavior
url https://arxiv.org/abs/2505.07865