Transcoder-based Circuit Analysis for Interpretable Single-Cell Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hosokawa, Sosuke, Kawakami, Toshiharu, Kodera, Satoshi, Ito, Masamichi, Takeda, Norihiko
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909795190046720
author Hosokawa, Sosuke
Kawakami, Toshiharu
Kodera, Satoshi
Ito, Masamichi
Takeda, Norihiko
author_facet Hosokawa, Sosuke
Kawakami, Toshiharu
Kodera, Satoshi
Ito, Masamichi
Takeda, Norihiko
contents Single-cell foundation models (scFMs) have demonstrated state-of-the-art performance on various tasks, such as cell-type annotation and perturbation response prediction, by learning gene regulatory networks from large-scale transcriptome data. However, a significant challenge remains: the decision-making processes of these models are less interpretable compared to traditional methods like differential gene expression analysis. Recently, transcoders have emerged as a promising approach for extracting interpretable decision circuits from large language models (LLMs). In this work, we train a transcoder on the cell2sentence (C2S) model, a state-of-the-art scFM. By leveraging the trained transcoder, we extract internal decision-making circuits from the C2S model. We demonstrate that the discovered circuits correspond to real-world biological mechanisms, confirming the potential of transcoders to uncover biologically plausible pathways within complex single-cell models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_14723
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Transcoder-based Circuit Analysis for Interpretable Single-Cell Foundation Models
Hosokawa, Sosuke
Kawakami, Toshiharu
Kodera, Satoshi
Ito, Masamichi
Takeda, Norihiko
Machine Learning
Single-cell foundation models (scFMs) have demonstrated state-of-the-art performance on various tasks, such as cell-type annotation and perturbation response prediction, by learning gene regulatory networks from large-scale transcriptome data. However, a significant challenge remains: the decision-making processes of these models are less interpretable compared to traditional methods like differential gene expression analysis. Recently, transcoders have emerged as a promising approach for extracting interpretable decision circuits from large language models (LLMs). In this work, we train a transcoder on the cell2sentence (C2S) model, a state-of-the-art scFM. By leveraging the trained transcoder, we extract internal decision-making circuits from the C2S model. We demonstrate that the discovered circuits correspond to real-world biological mechanisms, confirming the potential of transcoders to uncover biologically plausible pathways within complex single-cell models.
title Transcoder-based Circuit Analysis for Interpretable Single-Cell Foundation Models
topic Machine Learning
url https://arxiv.org/abs/2509.14723