TEDDY: A Family Of Foundation Models For Understanding Single Cell Biology

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chevalier, Alexis, Ghosh, Soumya, Awasthi, Urvi, Watkins, James, Bieniewska, Julia, Mitrea, Nichita, Kotova, Olga, Shkura, Kirill, Noble, Andrew, Steinbaugh, Michael, Sadashivaiah, Vijay, Dasoulas, George, Delile, Julien, Meier, Christoph, Zhukov, Leonid, Khalil, Iya, Mukherjee, Srayanta, Mueller, Judith
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910095257894912
author Chevalier, Alexis
Ghosh, Soumya
Awasthi, Urvi
Watkins, James
Bieniewska, Julia
Mitrea, Nichita
Kotova, Olga
Shkura, Kirill
Noble, Andrew
Steinbaugh, Michael
Sadashivaiah, Vijay
Dasoulas, George
Delile, Julien
Meier, Christoph
Zhukov, Leonid
Khalil, Iya
Mukherjee, Srayanta
Mueller, Judith
author_facet Chevalier, Alexis
Ghosh, Soumya
Awasthi, Urvi
Watkins, James
Bieniewska, Julia
Mitrea, Nichita
Kotova, Olga
Shkura, Kirill
Noble, Andrew
Steinbaugh, Michael
Sadashivaiah, Vijay
Dasoulas, George
Delile, Julien
Meier, Christoph
Zhukov, Leonid
Khalil, Iya
Mukherjee, Srayanta
Mueller, Judith
contents Understanding the biological mechanisms of disease is crucial for medicine, and in particular, for drug discovery. AI-powered analysis of genome-scale biological data holds great potential in this regard. The increasing availability of single-cell RNA sequencing data has enabled the development of large foundation models for disease biology. However, existing foundation models only modestly improve over task-specific models in downstream applications. Here, we explored two avenues for improving single-cell foundation models. First, we scaled the pre-training data to a diverse collection of 116 million cells, which is larger than those used by previous models. Second, we leveraged the availability of large-scale biological annotations as a form of supervision during pre-training. We trained the \model family of models comprising six transformer-based state-of-the-art single-cell foundation models with 70 million, 160 million, and 400 million parameters. We vetted our models on several downstream evaluation tasks, including identifying the underlying disease state of held-out donors not seen during training, distinguishing between diseased and healthy cells for disease conditions and donors not seen during training, and probing the learned representations for known biology. Our models showed substantial improvement over existing works, and scaling experiments showed that performance improved predictably with both data volume and parameter count.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03485
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TEDDY: A Family Of Foundation Models For Understanding Single Cell Biology
Chevalier, Alexis
Ghosh, Soumya
Awasthi, Urvi
Watkins, James
Bieniewska, Julia
Mitrea, Nichita
Kotova, Olga
Shkura, Kirill
Noble, Andrew
Steinbaugh, Michael
Sadashivaiah, Vijay
Dasoulas, George
Delile, Julien
Meier, Christoph
Zhukov, Leonid
Khalil, Iya
Mukherjee, Srayanta
Mueller, Judith
Machine Learning
Quantitative Methods
Understanding the biological mechanisms of disease is crucial for medicine, and in particular, for drug discovery. AI-powered analysis of genome-scale biological data holds great potential in this regard. The increasing availability of single-cell RNA sequencing data has enabled the development of large foundation models for disease biology. However, existing foundation models only modestly improve over task-specific models in downstream applications. Here, we explored two avenues for improving single-cell foundation models. First, we scaled the pre-training data to a diverse collection of 116 million cells, which is larger than those used by previous models. Second, we leveraged the availability of large-scale biological annotations as a form of supervision during pre-training. We trained the \model family of models comprising six transformer-based state-of-the-art single-cell foundation models with 70 million, 160 million, and 400 million parameters. We vetted our models on several downstream evaluation tasks, including identifying the underlying disease state of held-out donors not seen during training, distinguishing between diseased and healthy cells for disease conditions and donors not seen during training, and probing the learned representations for known biology. Our models showed substantial improvement over existing works, and scaling experiments showed that performance improved predictably with both data volume and parameter count.
title TEDDY: A Family Of Foundation Models For Understanding Single Cell Biology
topic Machine Learning
Quantitative Methods
url https://arxiv.org/abs/2503.03485