Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tang, Wenhao, Qin, Rong, Fang, Heng, Zhou, Fengtao, Chen, Hao, Li, Xiang, Cheng, Ming-Ming
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917034985521152
author Tang, Wenhao
Qin, Rong
Fang, Heng
Zhou, Fengtao
Chen, Hao
Li, Xiang
Cheng, Ming-Ming
author_facet Tang, Wenhao
Qin, Rong
Fang, Heng
Zhou, Fengtao
Chen, Hao
Li, Xiang
Cheng, Ming-Ming
contents Pre-trained encoders for offline feature extraction followed by multiple instance learning (MIL) aggregators have become the dominant paradigm in computational pathology (CPath), benefiting cancer diagnosis and prognosis. However, performance limitations arise from the absence of encoder fine-tuning for downstream tasks and disjoint optimization with MIL. While slide-level supervised end-to-end (E2E) learning is an intuitive solution to this issue, it faces challenges such as high computational demands and suboptimal results. These limitations motivate us to revisit E2E learning. We argue that prior work neglects inherent E2E optimization challenges, leading to performance disparities compared to traditional two-stage methods. In this paper, we pioneer the elucidation of optimization challenge caused by sparse-attention MIL and propose a novel MIL called ABMILX. It mitigates this problem through global correlation-based attention refinement and multi-head mechanisms. With the efficient multi-scale random patch sampling strategy, an E2E trained ResNet with ABMILX surpasses SOTA foundation models under the two-stage paradigm across multiple challenging benchmarks, while remaining computationally efficient (<10 RTX3090 hours). We show the potential of E2E learning in CPath and calls for greater research focus in this area. The code is https://github.com/DearCaat/E2E-WSI-ABMILX.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02408
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology
Tang, Wenhao
Qin, Rong
Fang, Heng
Zhou, Fengtao
Chen, Hao
Li, Xiang
Cheng, Ming-Ming
Computer Vision and Pattern Recognition
Pre-trained encoders for offline feature extraction followed by multiple instance learning (MIL) aggregators have become the dominant paradigm in computational pathology (CPath), benefiting cancer diagnosis and prognosis. However, performance limitations arise from the absence of encoder fine-tuning for downstream tasks and disjoint optimization with MIL. While slide-level supervised end-to-end (E2E) learning is an intuitive solution to this issue, it faces challenges such as high computational demands and suboptimal results. These limitations motivate us to revisit E2E learning. We argue that prior work neglects inherent E2E optimization challenges, leading to performance disparities compared to traditional two-stage methods. In this paper, we pioneer the elucidation of optimization challenge caused by sparse-attention MIL and propose a novel MIL called ABMILX. It mitigates this problem through global correlation-based attention refinement and multi-head mechanisms. With the efficient multi-scale random patch sampling strategy, an E2E trained ResNet with ABMILX surpasses SOTA foundation models under the two-stage paradigm across multiple challenging benchmarks, while remaining computationally efficient (<10 RTX3090 hours). We show the potential of E2E learning in CPath and calls for greater research focus in this area. The code is https://github.com/DearCaat/E2E-WSI-ABMILX.
title Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.02408