SemaMIL: Semantic-Aware Multiple Instance Learning with Retrieval-Guided State Space Modeling for Whole Slide Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gan, Lubin, Wu, Xiaoman, Zhang, Jing, Wang, Zhifeng, Qu, Linhao, Wu, Siying, Sun, Xiaoyan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918149592449024
author Gan, Lubin
Wu, Xiaoman
Zhang, Jing
Wang, Zhifeng
Qu, Linhao
Wu, Siying
Sun, Xiaoyan
author_facet Gan, Lubin
Wu, Xiaoman
Zhang, Jing
Wang, Zhifeng
Qu, Linhao
Wu, Siying
Sun, Xiaoyan
contents Multiple instance learning (MIL) has become the leading approach for extracting discriminative features from whole slide images (WSIs) in computational pathology. Attention-based MIL methods can identify key patches but tend to overlook contextual relationships. Transformer models are able to model interactions but require quadratic computational cost and are prone to overfitting. State space models (SSMs) offer linear complexity, yet shuffling patch order disrupts histological meaning and reduces interpretability. In this work, we introduce SemaMIL, which integrates Semantic Reordering (SR), an adaptive method that clusters and arranges semantically similar patches in sequence through a reversible permutation, with a Semantic-guided Retrieval State Space Module (SRSM) that chooses a representative subset of queries to adjust state space parameters for improved global modeling. Evaluation on four WSI subtype datasets shows that, compared to strong baselines, SemaMIL achieves state-of-the-art accuracy with fewer FLOPs and parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2509_00442
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SemaMIL: Semantic-Aware Multiple Instance Learning with Retrieval-Guided State Space Modeling for Whole Slide Images
Gan, Lubin
Wu, Xiaoman
Zhang, Jing
Wang, Zhifeng
Qu, Linhao
Wu, Siying
Sun, Xiaoyan
Computer Vision and Pattern Recognition
Multiple instance learning (MIL) has become the leading approach for extracting discriminative features from whole slide images (WSIs) in computational pathology. Attention-based MIL methods can identify key patches but tend to overlook contextual relationships. Transformer models are able to model interactions but require quadratic computational cost and are prone to overfitting. State space models (SSMs) offer linear complexity, yet shuffling patch order disrupts histological meaning and reduces interpretability. In this work, we introduce SemaMIL, which integrates Semantic Reordering (SR), an adaptive method that clusters and arranges semantically similar patches in sequence through a reversible permutation, with a Semantic-guided Retrieval State Space Module (SRSM) that chooses a representative subset of queries to adjust state space parameters for improved global modeling. Evaluation on four WSI subtype datasets shows that, compared to strong baselines, SemaMIL achieves state-of-the-art accuracy with fewer FLOPs and parameters.
title SemaMIL: Semantic-Aware Multiple Instance Learning with Retrieval-Guided State Space Modeling for Whole Slide Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.00442