FocusNet: Transformer-enhanced Polyp Segmentation with Local and Pooling Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Jun, Santosh, KC, Nayak, Deepak Rajan, de Lange, Thomas, Varkey, Jonas, Berzin, Tyler, Jha, Debesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916696195858432
author Zeng, Jun
Santosh, KC
Nayak, Deepak Rajan
de Lange, Thomas
Varkey, Jonas
Berzin, Tyler
Jha, Debesh
author_facet Zeng, Jun
Santosh, KC
Nayak, Deepak Rajan
de Lange, Thomas
Varkey, Jonas
Berzin, Tyler
Jha, Debesh
contents Colonoscopy is vital in the early diagnosis of colorectal polyps. Regular screenings can effectively prevent benign polyps from progressing to CRC. While deep learning has made impressive strides in polyp segmentation, most existing models are trained on single-modality and single-center data, making them less effective in real-world clinical environments. To overcome these limitations, we propose FocusNet, a Transformer-enhanced focus attention network designed to improve polyp segmentation. FocusNet incorporates three essential modules: the Cross-semantic Interaction Decoder Module (CIDM) for generating coarse segmentation maps, the Detail Enhancement Module (DEM) for refining shallow features, and the Focus Attention Module (FAM), to balance local detail and global context through local and pooling attention mechanisms. We evaluate our model on PolypDB, a newly introduced dataset with multi-modality and multi-center data for building more reliable segmentation methods. Extensive experiments showed that FocusNet consistently outperforms existing state-of-the-art approaches with a high dice coefficients of 82.47% on the BLI modality, 88.46% on FICE, 92.04% on LCI, 82.09% on the NBI and 93.42% on WLI modality, demonstrating its accuracy and robustness across five different modalities. The source code for FocusNet is available at https://github.com/JunZengz/FocusNet.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13597
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FocusNet: Transformer-enhanced Polyp Segmentation with Local and Pooling Attention
Zeng, Jun
Santosh, KC
Nayak, Deepak Rajan
de Lange, Thomas
Varkey, Jonas
Berzin, Tyler
Jha, Debesh
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Colonoscopy is vital in the early diagnosis of colorectal polyps. Regular screenings can effectively prevent benign polyps from progressing to CRC. While deep learning has made impressive strides in polyp segmentation, most existing models are trained on single-modality and single-center data, making them less effective in real-world clinical environments. To overcome these limitations, we propose FocusNet, a Transformer-enhanced focus attention network designed to improve polyp segmentation. FocusNet incorporates three essential modules: the Cross-semantic Interaction Decoder Module (CIDM) for generating coarse segmentation maps, the Detail Enhancement Module (DEM) for refining shallow features, and the Focus Attention Module (FAM), to balance local detail and global context through local and pooling attention mechanisms. We evaluate our model on PolypDB, a newly introduced dataset with multi-modality and multi-center data for building more reliable segmentation methods. Extensive experiments showed that FocusNet consistently outperforms existing state-of-the-art approaches with a high dice coefficients of 82.47% on the BLI modality, 88.46% on FICE, 92.04% on LCI, 82.09% on the NBI and 93.42% on WLI modality, demonstrating its accuracy and robustness across five different modalities. The source code for FocusNet is available at https://github.com/JunZengz/FocusNet.
title FocusNet: Transformer-enhanced Polyp Segmentation with Local and Pooling Attention
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.13597