TICON: A Slide-Level Tile Contextualizer for Histopathology Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Belagali, Varun, Kapse, Saarthak, Marza, Pierre, Das, Srijan, Li, Zilinghan, Boutaj, Sofiène, Pati, Pushpak, Yellapragada, Srikar, Nandi, Tarak Nath, Madduri, Ravi K, Saltz, Joel, Prasanna, Prateek, Christodoulidis, Stergios, Vakalopoulou, Maria, Samaras, Dimitris
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912788917518336
author Belagali, Varun
Kapse, Saarthak
Marza, Pierre
Das, Srijan
Li, Zilinghan
Boutaj, Sofiène
Pati, Pushpak
Yellapragada, Srikar
Nandi, Tarak Nath
Madduri, Ravi K
Saltz, Joel
Prasanna, Prateek
Christodoulidis, Stergios
Vakalopoulou, Maria
Samaras, Dimitris
author_facet Belagali, Varun
Kapse, Saarthak
Marza, Pierre
Das, Srijan
Li, Zilinghan
Boutaj, Sofiène
Pati, Pushpak
Yellapragada, Srikar
Nandi, Tarak Nath
Madduri, Ravi K
Saltz, Joel
Prasanna, Prateek
Christodoulidis, Stergios
Vakalopoulou, Maria
Samaras, Dimitris
contents The interpretation of small tiles in large whole slide images (WSI) often needs a larger image context. We introduce TICON, a transformer-based tile representation contextualizer that produces rich, contextualized embeddings for ''any'' application in computational pathology. Standard tile encoder-based pipelines, which extract embeddings of tiles stripped from their context, fail to model the rich slide-level information essential for both local and global tasks. Furthermore, different tile-encoders excel at different downstream tasks. Therefore, a unified model is needed to contextualize embeddings derived from ''any'' tile-level foundation model. TICON addresses this need with a single, shared encoder, pretrained using a masked modeling objective to simultaneously unify and contextualize representations from diverse tile-level pathology foundation models. Our experiments demonstrate that TICON-contextualized embeddings significantly improve performance across many different tasks, establishing new state-of-the-art results on tile-level benchmarks (i.e., HEST-Bench, THUNDER, CATCH) and slide-level benchmarks (i.e., Patho-Bench). Finally, we pretrain an aggregator on TICON to form a slide-level foundation model, using only 11K WSIs, outperforming SoTA slide-level foundation models pretrained with up to 350K WSIs.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21331
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TICON: A Slide-Level Tile Contextualizer for Histopathology Representation Learning
Belagali, Varun
Kapse, Saarthak
Marza, Pierre
Das, Srijan
Li, Zilinghan
Boutaj, Sofiène
Pati, Pushpak
Yellapragada, Srikar
Nandi, Tarak Nath
Madduri, Ravi K
Saltz, Joel
Prasanna, Prateek
Christodoulidis, Stergios
Vakalopoulou, Maria
Samaras, Dimitris
Computer Vision and Pattern Recognition
The interpretation of small tiles in large whole slide images (WSI) often needs a larger image context. We introduce TICON, a transformer-based tile representation contextualizer that produces rich, contextualized embeddings for ''any'' application in computational pathology. Standard tile encoder-based pipelines, which extract embeddings of tiles stripped from their context, fail to model the rich slide-level information essential for both local and global tasks. Furthermore, different tile-encoders excel at different downstream tasks. Therefore, a unified model is needed to contextualize embeddings derived from ''any'' tile-level foundation model. TICON addresses this need with a single, shared encoder, pretrained using a masked modeling objective to simultaneously unify and contextualize representations from diverse tile-level pathology foundation models. Our experiments demonstrate that TICON-contextualized embeddings significantly improve performance across many different tasks, establishing new state-of-the-art results on tile-level benchmarks (i.e., HEST-Bench, THUNDER, CATCH) and slide-level benchmarks (i.e., Patho-Bench). Finally, we pretrain an aggregator on TICON to form a slide-level foundation model, using only 11K WSIs, outperforming SoTA slide-level foundation models pretrained with up to 350K WSIs.
title TICON: A Slide-Level Tile Contextualizer for Histopathology Representation Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.21331