HypoGeneAgent: A Hypothesis Language Agent for Gene-Set Cluster Resolution Selection Using Perturb-seq Datasets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Ying, Ge, Xing-Yue Monica, Waterman, Aaron Archer, Biancalani, Tommaso, Richmond, David, Pandit, Yogesh, Singh, Avtar, Littman, Russell, Liu, Jin, Huetter, Jan-Christian, Ermakov, Vladimir
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915491329605632
author Yuan, Ying
Ge, Xing-Yue Monica
Waterman, Aaron Archer
Biancalani, Tommaso
Richmond, David
Pandit, Yogesh
Singh, Avtar
Littman, Russell
Liu, Jin
Huetter, Jan-Christian
Ermakov, Vladimir
author_facet Yuan, Ying
Ge, Xing-Yue Monica
Waterman, Aaron Archer
Biancalani, Tommaso
Richmond, David
Pandit, Yogesh
Singh, Avtar
Littman, Russell
Liu, Jin
Huetter, Jan-Christian
Ermakov, Vladimir
contents Large-scale single-cell and Perturb-seq investigations routinely involve clustering cells and subsequently annotating each cluster with Gene-Ontology (GO) terms to elucidate the underlying biological programs. However, both stages, resolution selection and functional annotation, are inherently subjective, relying on heuristics and expert curation. We present HYPOGENEAGENT, a large language model (LLM)-driven framework, transforming cluster annotation into a quantitatively optimizable task. Initially, an LLM functioning as a gene-set analyst analyzes the content of each gene program or perturbation module and generates a ranked list of GO-based hypotheses, accompanied by calibrated confidence scores. Subsequently, we embed every predicted description with a sentence-embedding model, compute pair-wise cosine similarities, and let the agent referee panel score (i) the internal consistency of the predictions, high average similarity within the same cluster, termed intra-cluster agreement (ii) their external distinctiveness, low similarity between clusters, termed inter-cluster separation. These two quantities are combined to produce an agent-derived resolution score, which is maximized when clusters exhibit simultaneous coherence and mutual exclusivity. When applied to a public K562 CRISPRi Perturb-seq dataset as a preliminary test, our Resolution Score selects clustering granularities that exhibit alignment with known pathway compared to classical metrics such silhouette score, modularity score for gene functional enrichment summary. These findings establish LLM agents as objective adjudicators of cluster resolution and functional annotation, thereby paving the way for fully automated, context-aware interpretation pipelines in single-cell multi-omics studies.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09740
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HypoGeneAgent: A Hypothesis Language Agent for Gene-Set Cluster Resolution Selection Using Perturb-seq Datasets
Yuan, Ying
Ge, Xing-Yue Monica
Waterman, Aaron Archer
Biancalani, Tommaso
Richmond, David
Pandit, Yogesh
Singh, Avtar
Littman, Russell
Liu, Jin
Huetter, Jan-Christian
Ermakov, Vladimir
Quantitative Methods
Artificial Intelligence
Computation and Language
Machine Learning
Large-scale single-cell and Perturb-seq investigations routinely involve clustering cells and subsequently annotating each cluster with Gene-Ontology (GO) terms to elucidate the underlying biological programs. However, both stages, resolution selection and functional annotation, are inherently subjective, relying on heuristics and expert curation. We present HYPOGENEAGENT, a large language model (LLM)-driven framework, transforming cluster annotation into a quantitatively optimizable task. Initially, an LLM functioning as a gene-set analyst analyzes the content of each gene program or perturbation module and generates a ranked list of GO-based hypotheses, accompanied by calibrated confidence scores. Subsequently, we embed every predicted description with a sentence-embedding model, compute pair-wise cosine similarities, and let the agent referee panel score (i) the internal consistency of the predictions, high average similarity within the same cluster, termed intra-cluster agreement (ii) their external distinctiveness, low similarity between clusters, termed inter-cluster separation. These two quantities are combined to produce an agent-derived resolution score, which is maximized when clusters exhibit simultaneous coherence and mutual exclusivity. When applied to a public K562 CRISPRi Perturb-seq dataset as a preliminary test, our Resolution Score selects clustering granularities that exhibit alignment with known pathway compared to classical metrics such silhouette score, modularity score for gene functional enrichment summary. These findings establish LLM agents as objective adjudicators of cluster resolution and functional annotation, thereby paving the way for fully automated, context-aware interpretation pipelines in single-cell multi-omics studies.
title HypoGeneAgent: A Hypothesis Language Agent for Gene-Set Cluster Resolution Selection Using Perturb-seq Datasets
topic Quantitative Methods
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2509.09740