ColonAdapter: Geometry Estimation Through Foundation Model Adaptation for Colonoscopy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Zhiyi, Wang, Yifu, Cheng, Xuelian, Ge, Zongyuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918221083312128
author Jiang, Zhiyi
Wang, Yifu
Cheng, Xuelian
Ge, Zongyuan
author_facet Jiang, Zhiyi
Wang, Yifu
Cheng, Xuelian
Ge, Zongyuan
contents Estimating 3D geometry from monocular colonoscopy images is challenging due to non-Lambertian surfaces, moving light sources, and large textureless regions. While recent 3D geometric foundation models eliminate the need for multi-stage pipelines, their performance deteriorates in clinical scenes. These models are primarily trained on natural scene datasets and struggle with specularity and homogeneous textures typical in colonoscopy, leading to inaccurate geometry estimation. In this paper, we present ColonAdapter, a self-supervised fine-tuning framework that adapts geometric foundation models for colonoscopy geometry estimation. Our method leverages pretrained geometric priors while tailoring them to clinical data. To improve performance in low-texture regions and ensure scale consistency, we introduce a Detail Restoration Module (DRM) and a geometry consistency loss. Furthermore, a confidence-weighted photometric loss enhances training stability in clinical environments. Experiments on both synthetic and real datasets demonstrate that our approach achieves state-of-the-art performance in camera pose estimation, monocular depth prediction, and dense 3D point map reconstruction, without requiring ground-truth intrinsic parameters.
format Preprint
id arxiv_https___arxiv_org_abs_2511_22250
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ColonAdapter: Geometry Estimation Through Foundation Model Adaptation for Colonoscopy
Jiang, Zhiyi
Wang, Yifu
Cheng, Xuelian
Ge, Zongyuan
Image and Video Processing
Computer Vision and Pattern Recognition
Estimating 3D geometry from monocular colonoscopy images is challenging due to non-Lambertian surfaces, moving light sources, and large textureless regions. While recent 3D geometric foundation models eliminate the need for multi-stage pipelines, their performance deteriorates in clinical scenes. These models are primarily trained on natural scene datasets and struggle with specularity and homogeneous textures typical in colonoscopy, leading to inaccurate geometry estimation. In this paper, we present ColonAdapter, a self-supervised fine-tuning framework that adapts geometric foundation models for colonoscopy geometry estimation. Our method leverages pretrained geometric priors while tailoring them to clinical data. To improve performance in low-texture regions and ensure scale consistency, we introduce a Detail Restoration Module (DRM) and a geometry consistency loss. Furthermore, a confidence-weighted photometric loss enhances training stability in clinical environments. Experiments on both synthetic and real datasets demonstrate that our approach achieves state-of-the-art performance in camera pose estimation, monocular depth prediction, and dense 3D point map reconstruction, without requiring ground-truth intrinsic parameters.
title ColonAdapter: Geometry Estimation Through Foundation Model Adaptation for Colonoscopy
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.22250