EndoCaver: Handling Fog, Blur and Glare in Endoscopic Images via Joint Deblurring-Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Zhuoyu, Ou, Wenhui, Tan, Pei-Sze, Yang, Jiayan, Fang, Wenqi, Wang, Zheng, Phan, Raphaël C. -W.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911689964781568
author Wu, Zhuoyu
Ou, Wenhui
Tan, Pei-Sze
Yang, Jiayan
Fang, Wenqi
Wang, Zheng
Phan, Raphaël C. -W.
author_facet Wu, Zhuoyu
Ou, Wenhui
Tan, Pei-Sze
Yang, Jiayan
Fang, Wenqi
Wang, Zheng
Phan, Raphaël C. -W.
contents Endoscopic image analysis is vital for colorectal cancer screening, yet real-world conditions often suffer from lens fogging, motion blur, and specular highlights, which severely compromise automated polyp detection. We propose EndoCaver, a lightweight transformer with a unidirectional-guided dual-decoder architecture, enabling joint multi-task capability for image deblurring and segmentation while significantly reducing computational complexity and model parameters. Specifically, it integrates a Global Attention Module (GAM) for cross-scale aggregation, a Deblurring-Segmentation Aligner (DSA) to transfer restoration cues, and a cosine-based scheduler (LoCoS) for stable multi-task optimisation. Experiments on the Kvasir-SEG dataset show that EndoCaver achieves 0.922 Dice on clean data and 0.889 under severe image degradation, surpassing state-of-the-art methods while reducing model parameters by 90%. These results demonstrate its efficiency and robustness, making it well-suited for on-device clinical deployment. Code is available at https://github.com/ReaganWu/EndoCaver.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22537
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle EndoCaver: Handling Fog, Blur and Glare in Endoscopic Images via Joint Deblurring-Segmentation
Wu, Zhuoyu
Ou, Wenhui
Tan, Pei-Sze
Yang, Jiayan
Fang, Wenqi
Wang, Zheng
Phan, Raphaël C. -W.
Image and Video Processing
Computer Vision and Pattern Recognition
Endoscopic image analysis is vital for colorectal cancer screening, yet real-world conditions often suffer from lens fogging, motion blur, and specular highlights, which severely compromise automated polyp detection. We propose EndoCaver, a lightweight transformer with a unidirectional-guided dual-decoder architecture, enabling joint multi-task capability for image deblurring and segmentation while significantly reducing computational complexity and model parameters. Specifically, it integrates a Global Attention Module (GAM) for cross-scale aggregation, a Deblurring-Segmentation Aligner (DSA) to transfer restoration cues, and a cosine-based scheduler (LoCoS) for stable multi-task optimisation. Experiments on the Kvasir-SEG dataset show that EndoCaver achieves 0.922 Dice on clean data and 0.889 under severe image degradation, surpassing state-of-the-art methods while reducing model parameters by 90%. These results demonstrate its efficiency and robustness, making it well-suited for on-device clinical deployment. Code is available at https://github.com/ReaganWu/EndoCaver.
title EndoCaver: Handling Fog, Blur and Glare in Endoscopic Images via Joint Deblurring-Segmentation
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.22537