MambaMIL+: Modeling Long-Term Contextual Patterns for Gigapixel Whole Slide Image

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zeng, Qian, Wang, Yihui, Yang, Shu, Xu, Yingxue, Zhou, Fengtao, Ma, Jiabo, Cai, Dejia, Zhang, Zhengyu, Qu, Lijuan, Wang, Yu, Liang, Li, Chen, Hao
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914210692202496
author Zeng, Qian
Wang, Yihui
Yang, Shu
Xu, Yingxue
Zhou, Fengtao
Ma, Jiabo
Cai, Dejia
Zhang, Zhengyu
Qu, Lijuan
Wang, Yu
Liang, Li
Chen, Hao
author_facet Zeng, Qian
Wang, Yihui
Yang, Shu
Xu, Yingxue
Zhou, Fengtao
Ma, Jiabo
Cai, Dejia
Zhang, Zhengyu
Qu, Lijuan
Wang, Yu
Liang, Li
Chen, Hao
contents Whole-slide images (WSIs) are an important data modality in computational pathology, yet their gigapixel resolution and lack of fine-grained annotations challenge conventional deep learning models. Multiple instance learning (MIL) offers a solution by treating each WSI as a bag of patch-level instances, but effectively modeling ultra-long sequences with rich spatial context remains difficult. Recently, Mamba has emerged as a promising alternative for long sequence learning, scaling linearly to thousands of tokens. However, despite its efficiency, it still suffers from limited spatial context modeling and memory decay, constraining its effectiveness to WSI analysis. To address these limitations, we propose MambaMIL+, a new MIL framework that explicitly integrates spatial context while maintaining long-range dependency modeling without memory forgetting. Specifically, MambaMIL+ introduces 1) overlapping scanning, which restructures the patch sequence to embed spatial continuity and instance correlations; 2) a selective stripe position encoder (S2PE) that encodes positional information while mitigating the biases of fixed scanning orders; and 3) a contextual token selection (CTS) mechanism, which leverages supervisory knowledge to dynamically enlarge the contextual memory for stable long-range modeling. Extensive experiments on 20 benchmarks across diagnostic classification, molecular prediction, and survival analysis demonstrate that MambaMIL+ consistently achieves state-of-the-art performance under three feature extractors (ResNet-50, PLIP, and CONCH), highlighting its effectiveness and robustness for large-scale computational pathology
format Preprint
id arxiv_https___arxiv_org_abs_2512_17726
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MambaMIL+: Modeling Long-Term Contextual Patterns for Gigapixel Whole Slide Image
Zeng, Qian
Wang, Yihui
Yang, Shu
Xu, Yingxue
Zhou, Fengtao
Ma, Jiabo
Cai, Dejia
Zhang, Zhengyu
Qu, Lijuan
Wang, Yu
Liang, Li
Chen, Hao
Computer Vision and Pattern Recognition
Whole-slide images (WSIs) are an important data modality in computational pathology, yet their gigapixel resolution and lack of fine-grained annotations challenge conventional deep learning models. Multiple instance learning (MIL) offers a solution by treating each WSI as a bag of patch-level instances, but effectively modeling ultra-long sequences with rich spatial context remains difficult. Recently, Mamba has emerged as a promising alternative for long sequence learning, scaling linearly to thousands of tokens. However, despite its efficiency, it still suffers from limited spatial context modeling and memory decay, constraining its effectiveness to WSI analysis. To address these limitations, we propose MambaMIL+, a new MIL framework that explicitly integrates spatial context while maintaining long-range dependency modeling without memory forgetting. Specifically, MambaMIL+ introduces 1) overlapping scanning, which restructures the patch sequence to embed spatial continuity and instance correlations; 2) a selective stripe position encoder (S2PE) that encodes positional information while mitigating the biases of fixed scanning orders; and 3) a contextual token selection (CTS) mechanism, which leverages supervisory knowledge to dynamically enlarge the contextual memory for stable long-range modeling. Extensive experiments on 20 benchmarks across diagnostic classification, molecular prediction, and survival analysis demonstrate that MambaMIL+ consistently achieves state-of-the-art performance under three feature extractors (ResNet-50, PLIP, and CONCH), highlighting its effectiveness and robustness for large-scale computational pathology
title MambaMIL+: Modeling Long-Term Contextual Patterns for Gigapixel Whole Slide Image
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.17726