scShapeBench: Discovering geometry from high dimensional scRNAseq data

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Steindl, Andrew J, Rocha, João Felipe, Di Bassinga, Brian Tshilengi, Warren, Zachary, Scicluna, Matthew, Córdova, César Miguel Valdez, Gupta, Shabarni, Torices, Leire, Neumann, Daniel, Mann, Timothy J., Gunawan, Ihuan, Bhaskar, Dhananjay, Lock, John G, Chaffer, Christine L, Wolf, Guy, Krishnaswamy, Smita
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917488421240832
author Steindl, Andrew J
Rocha, João Felipe
Di Bassinga, Brian Tshilengi
Warren, Zachary
Scicluna, Matthew
Córdova, César Miguel Valdez
Gupta, Shabarni
Torices, Leire
Neumann, Daniel
Mann, Timothy J.
Gunawan, Ihuan
Bhaskar, Dhananjay
Lock, John G
Chaffer, Christine L
Wolf, Guy
Krishnaswamy, Smita
author_facet Steindl, Andrew J
Rocha, João Felipe
Di Bassinga, Brian Tshilengi
Warren, Zachary
Scicluna, Matthew
Córdova, César Miguel Valdez
Gupta, Shabarni
Torices, Leire
Neumann, Daniel
Mann, Timothy J.
Gunawan, Ihuan
Bhaskar, Dhananjay
Lock, John G
Chaffer, Christine L
Wolf, Guy
Krishnaswamy, Smita
contents High-dimensional point cloud data arise across many scientific domains, especially single-cell biology. The shapes or topologies of these datasets determine the types of information that can be extracted. For example, clustered data supports cell-type identification, trajectory structures support transition analysis, and archetypal structures capture continua of cellular behaviors. Existing analysis pipelines often assume a specific shape. The standard Seurat pipeline combines UMAP visualization with Louvain clustering and therefore assumes clustered data, while tools such as Monocle and SPADE assume tree-like structures, and flow-based models such as MIOFlow and Conditional Flow Matching target trajectories. Choosing which pipeline to apply is therefore often left to bioinformaticians who visually inspect datasets before selecting an analysis strategy. With the rise of agentic AI scientists, automating shape detection is increasingly important for selecting downstream analysis pipelines. To address this problem, we introduce scShapeBench, a benchmark dataset for shape detection containing both synthetic and expert-annotated single-cell datasets. Synthetic datasets are sampled from ground-truth skeleton graphs with controlled variance. Real single-cell datasets are curated from diverse sources and annotated by experts into four categories: clusters, single trajectory, multi-branching, and archetypal. We additionally introduce scReebTower, a baseline method that uses diffusion geometry to extract Reeb graphs and connect visualization with pipeline selection. We provide topology-aware evaluation metrics and compare scReebTower against PAGA and Mapper on synthetic and real data. Our results indicate that scReebTower outperforms existing baselines. Overall, our contributions span benchmarks, evaluation metrics, and a baseline for automated shape detection in single-cell data.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12662
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle scShapeBench: Discovering geometry from high dimensional scRNAseq data
Steindl, Andrew J
Rocha, João Felipe
Di Bassinga, Brian Tshilengi
Warren, Zachary
Scicluna, Matthew
Córdova, César Miguel Valdez
Gupta, Shabarni
Torices, Leire
Neumann, Daniel
Mann, Timothy J.
Gunawan, Ihuan
Bhaskar, Dhananjay
Lock, John G
Chaffer, Christine L
Wolf, Guy
Krishnaswamy, Smita
Machine Learning
Genomics
High-dimensional point cloud data arise across many scientific domains, especially single-cell biology. The shapes or topologies of these datasets determine the types of information that can be extracted. For example, clustered data supports cell-type identification, trajectory structures support transition analysis, and archetypal structures capture continua of cellular behaviors. Existing analysis pipelines often assume a specific shape. The standard Seurat pipeline combines UMAP visualization with Louvain clustering and therefore assumes clustered data, while tools such as Monocle and SPADE assume tree-like structures, and flow-based models such as MIOFlow and Conditional Flow Matching target trajectories. Choosing which pipeline to apply is therefore often left to bioinformaticians who visually inspect datasets before selecting an analysis strategy. With the rise of agentic AI scientists, automating shape detection is increasingly important for selecting downstream analysis pipelines. To address this problem, we introduce scShapeBench, a benchmark dataset for shape detection containing both synthetic and expert-annotated single-cell datasets. Synthetic datasets are sampled from ground-truth skeleton graphs with controlled variance. Real single-cell datasets are curated from diverse sources and annotated by experts into four categories: clusters, single trajectory, multi-branching, and archetypal. We additionally introduce scReebTower, a baseline method that uses diffusion geometry to extract Reeb graphs and connect visualization with pipeline selection. We provide topology-aware evaluation metrics and compare scReebTower against PAGA and Mapper on synthetic and real data. Our results indicate that scReebTower outperforms existing baselines. Overall, our contributions span benchmarks, evaluation metrics, and a baseline for automated shape detection in single-cell data.
title scShapeBench: Discovering geometry from high dimensional scRNAseq data
topic Machine Learning
Genomics
url https://arxiv.org/abs/2605.12662