Dino-NestedUNet: Unlocking Foundation Vision Encoders for Pathology Tumor Bulk Segmentation via Dense Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Tianyang, Su, Ziyu, Akbar, Abdul Rehman, Sajjad, Usama, Afzaal, Usman, Gokhale, Lina, Rabolli, Charles, Chen, Wei, Parwani, Anil, Niazi, Muhammad Khalid Khan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914525873176576
author Wang, Tianyang
Su, Ziyu
Akbar, Abdul Rehman
Sajjad, Usama
Afzaal, Usman
Gokhale, Lina
Rabolli, Charles
Chen, Wei
Parwani, Anil
Niazi, Muhammad Khalid Khan
author_facet Wang, Tianyang
Su, Ziyu
Akbar, Abdul Rehman
Sajjad, Usama
Afzaal, Usman
Gokhale, Lina
Rabolli, Charles
Chen, Wei
Parwani, Anil
Niazi, Muhammad Khalid Khan
contents Vision foundation models (VFMs), such as DINOv3, provide rich semantic representations that are promising for computational pathology. However, many current adaptations pair frozen VFMs with lightweight decoders, creating a capacity mismatch that often limits boundary fidelity for infiltrative tumor bulk segmentation. This paper presents Dino-NestedUNet, a framework that couples a pre-trained DINOv3 encoder with a Nested Dense Decoder. Instead of sparse skip connections and linear upsampling, the proposed decoder forms a dense grid of intermediate pathways to enable continuous feature reuse and multi-scale recalibration, aligning high-level semantics with low-level morphological textures during reconstruction. We evaluate Dino-NestedUNet on three histopathology cohorts (multi-center CHTN, institutional OSU, and CAMELYON16) and observe consistent improvements over UNet++ and standard Dino-UNet variants, particularly under cross-domain shift. To further assess external generalization, we perform zero-shot evaluation by training on CHTN and directly testing on unseen TIGER WSIBULK and OSU CRC cohorts without fine-tuning. These results suggest that dense decoding is a key ingredient for unlocking foundation encoders in boundary-sensitive pathology segmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2605_00894
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Dino-NestedUNet: Unlocking Foundation Vision Encoders for Pathology Tumor Bulk Segmentation via Dense Decoding
Wang, Tianyang
Su, Ziyu
Akbar, Abdul Rehman
Sajjad, Usama
Afzaal, Usman
Gokhale, Lina
Rabolli, Charles
Chen, Wei
Parwani, Anil
Niazi, Muhammad Khalid Khan
Computer Vision and Pattern Recognition
Vision foundation models (VFMs), such as DINOv3, provide rich semantic representations that are promising for computational pathology. However, many current adaptations pair frozen VFMs with lightweight decoders, creating a capacity mismatch that often limits boundary fidelity for infiltrative tumor bulk segmentation. This paper presents Dino-NestedUNet, a framework that couples a pre-trained DINOv3 encoder with a Nested Dense Decoder. Instead of sparse skip connections and linear upsampling, the proposed decoder forms a dense grid of intermediate pathways to enable continuous feature reuse and multi-scale recalibration, aligning high-level semantics with low-level morphological textures during reconstruction. We evaluate Dino-NestedUNet on three histopathology cohorts (multi-center CHTN, institutional OSU, and CAMELYON16) and observe consistent improvements over UNet++ and standard Dino-UNet variants, particularly under cross-domain shift. To further assess external generalization, we perform zero-shot evaluation by training on CHTN and directly testing on unseen TIGER WSIBULK and OSU CRC cohorts without fine-tuning. These results suggest that dense decoding is a key ingredient for unlocking foundation encoders in boundary-sensitive pathology segmentation.
title Dino-NestedUNet: Unlocking Foundation Vision Encoders for Pathology Tumor Bulk Segmentation via Dense Decoding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.00894