Learning biologically relevant features in a pathology foundation model using sparse autoencoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Le, Nhat Minh, Shen, Ciyue, Patel, Neel, Shah, Chintan, Sanghavi, Darpan, Martin, Blake, Eng, Alfred, Shenker, Daniel, Padigela, Harshith, Biju, Raymond, Javed, Syed Ashar, Hipp, Jennifer, Abel, John, Pokkalla, Harsha, Grullon, Sean, Juyal, Dinkar
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913614439383040
author Le, Nhat Minh
Shen, Ciyue
Patel, Neel
Shah, Chintan
Sanghavi, Darpan
Martin, Blake
Eng, Alfred
Shenker, Daniel
Padigela, Harshith
Biju, Raymond
Javed, Syed Ashar
Hipp, Jennifer
Abel, John
Pokkalla, Harsha
Grullon, Sean
Juyal, Dinkar
author_facet Le, Nhat Minh
Shen, Ciyue
Patel, Neel
Shah, Chintan
Sanghavi, Darpan
Martin, Blake
Eng, Alfred
Shenker, Daniel
Padigela, Harshith
Biju, Raymond
Javed, Syed Ashar
Hipp, Jennifer
Abel, John
Pokkalla, Harsha
Grullon, Sean
Juyal, Dinkar
contents Pathology plays an important role in disease diagnosis, treatment decision-making and drug development. Previous works on interpretability for machine learning models on pathology images have revolved around methods such as attention value visualization and deriving human-interpretable features from model heatmaps. Mechanistic interpretability is an emerging area of model interpretability that focuses on reverse-engineering neural networks. Sparse Autoencoders (SAEs) have emerged as a promising direction in terms of extracting monosemantic features from polysemantic model activations. In this work, we trained a Sparse Autoencoder on the embeddings of a pathology pretrained foundation model. We found that Sparse Autoencoder features represent interpretable and monosemantic biological concepts. In particular, individual SAE dimensions showed strong correlations with cell type counts such as plasma cells and lymphocytes. These biological representations were unique to the pathology pretrained model and were not found in a self-supervised model pretrained on natural images. We demonstrated that such biologically-grounded monosemantic representations evolved across the model's depth, and the pathology foundation model eventually gained robustness to non-biological factors such as scanner type. The emergence of biologically relevant SAE features was generalizable to an out-of-domain dataset. Our work paves the way for further exploration around interpretable feature dimensions and their utility for medical and clinical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10785
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Learning biologically relevant features in a pathology foundation model using sparse autoencoders
Le, Nhat Minh
Shen, Ciyue
Patel, Neel
Shah, Chintan
Sanghavi, Darpan
Martin, Blake
Eng, Alfred
Shenker, Daniel
Padigela, Harshith
Biju, Raymond
Javed, Syed Ashar
Hipp, Jennifer
Abel, John
Pokkalla, Harsha
Grullon, Sean
Juyal, Dinkar
Image and Video Processing
Computer Vision and Pattern Recognition
Pathology plays an important role in disease diagnosis, treatment decision-making and drug development. Previous works on interpretability for machine learning models on pathology images have revolved around methods such as attention value visualization and deriving human-interpretable features from model heatmaps. Mechanistic interpretability is an emerging area of model interpretability that focuses on reverse-engineering neural networks. Sparse Autoencoders (SAEs) have emerged as a promising direction in terms of extracting monosemantic features from polysemantic model activations. In this work, we trained a Sparse Autoencoder on the embeddings of a pathology pretrained foundation model. We found that Sparse Autoencoder features represent interpretable and monosemantic biological concepts. In particular, individual SAE dimensions showed strong correlations with cell type counts such as plasma cells and lymphocytes. These biological representations were unique to the pathology pretrained model and were not found in a self-supervised model pretrained on natural images. We demonstrated that such biologically-grounded monosemantic representations evolved across the model's depth, and the pathology foundation model eventually gained robustness to non-biological factors such as scanner type. The emergence of biologically relevant SAE features was generalizable to an out-of-domain dataset. Our work paves the way for further exploration around interpretable feature dimensions and their utility for medical and clinical applications.
title Learning biologically relevant features in a pathology foundation model using sparse autoencoders
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.10785