Benchmarking histopathology foundation models in a multi-center dataset for skin cancer subtyping

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meseguer, Pablo, del Amor, Rocío, Naranjo, Valery
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911067349712896
author Meseguer, Pablo
del Amor, Rocío
Naranjo, Valery
author_facet Meseguer, Pablo
del Amor, Rocío
Naranjo, Valery
contents Pretraining on large-scale, in-domain datasets grants histopathology foundation models (FM) the ability to learn task-agnostic data representations, enhancing transfer learning on downstream tasks. In computational pathology, automated whole slide image analysis requires multiple instance learning (MIL) frameworks due to the gigapixel scale of the slides. The diversity among histopathology FMs has highlighted the need to design real-world challenges for evaluating their effectiveness. To bridge this gap, our work presents a novel benchmark for evaluating histopathology FMs as patch-level feature extractors within a MIL classification framework. For that purpose, we leverage the AI4SkIN dataset, a multi-center cohort encompassing slides with challenging cutaneous spindle cell neoplasm subtypes. We also define the Foundation Model - Silhouette Index (FM-SI), a novel metric to measure model consistency against distribution shifts. Our experimentation shows that extracting less biased features enhances classification performance, especially in similarity-based MIL classifiers.
format Preprint
id arxiv_https___arxiv_org_abs_2506_18668
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking histopathology foundation models in a multi-center dataset for skin cancer subtyping
Meseguer, Pablo
del Amor, Rocío
Naranjo, Valery
Computer Vision and Pattern Recognition
Artificial Intelligence
Pretraining on large-scale, in-domain datasets grants histopathology foundation models (FM) the ability to learn task-agnostic data representations, enhancing transfer learning on downstream tasks. In computational pathology, automated whole slide image analysis requires multiple instance learning (MIL) frameworks due to the gigapixel scale of the slides. The diversity among histopathology FMs has highlighted the need to design real-world challenges for evaluating their effectiveness. To bridge this gap, our work presents a novel benchmark for evaluating histopathology FMs as patch-level feature extractors within a MIL classification framework. For that purpose, we leverage the AI4SkIN dataset, a multi-center cohort encompassing slides with challenging cutaneous spindle cell neoplasm subtypes. We also define the Foundation Model - Silhouette Index (FM-SI), a novel metric to measure model consistency against distribution shifts. Our experimentation shows that extracting less biased features enhances classification performance, especially in similarity-based MIL classifiers.
title Benchmarking histopathology foundation models in a multi-center dataset for skin cancer subtyping
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2506.18668