ColonCrafter: A Depth Estimation Model for Colonoscopy Videos Using Diffusion Priors

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hardy, Romain, Berzin, Tyler, Rajpurkar, Pranav
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912590125334528
author Hardy, Romain
Berzin, Tyler
Rajpurkar, Pranav
author_facet Hardy, Romain
Berzin, Tyler
Rajpurkar, Pranav
contents Three-dimensional (3D) scene understanding in colonoscopy presents significant challenges that necessitate automated methods for accurate depth estimation. However, existing depth estimation models for endoscopy struggle with temporal consistency across video sequences, limiting their applicability for 3D reconstruction. We present ColonCrafter, a diffusion-based depth estimation model that generates temporally consistent depth maps from monocular colonoscopy videos. Our approach learns robust geometric priors from synthetic colonoscopy sequences to generate temporally consistent depth maps. We also introduce a style transfer technique that preserves geometric structure while adapting real clinical videos to match our synthetic training domain. ColonCrafter achieves state-of-the-art zero-shot performance on the C3VD dataset, outperforming both general-purpose and endoscopy-specific approaches. Although full trajectory 3D reconstruction remains a challenge, we demonstrate clinically relevant applications of ColonCrafter, including 3D point cloud generation and surface coverage assessment.
format Preprint
id arxiv_https___arxiv_org_abs_2509_13525
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ColonCrafter: A Depth Estimation Model for Colonoscopy Videos Using Diffusion Priors
Hardy, Romain
Berzin, Tyler
Rajpurkar, Pranav
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Three-dimensional (3D) scene understanding in colonoscopy presents significant challenges that necessitate automated methods for accurate depth estimation. However, existing depth estimation models for endoscopy struggle with temporal consistency across video sequences, limiting their applicability for 3D reconstruction. We present ColonCrafter, a diffusion-based depth estimation model that generates temporally consistent depth maps from monocular colonoscopy videos. Our approach learns robust geometric priors from synthetic colonoscopy sequences to generate temporally consistent depth maps. We also introduce a style transfer technique that preserves geometric structure while adapting real clinical videos to match our synthetic training domain. ColonCrafter achieves state-of-the-art zero-shot performance on the C3VD dataset, outperforming both general-purpose and endoscopy-specific approaches. Although full trajectory 3D reconstruction remains a challenge, we demonstrate clinically relevant applications of ColonCrafter, including 3D point cloud generation and surface coverage assessment.
title ColonCrafter: A Depth Estimation Model for Colonoscopy Videos Using Diffusion Priors
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.13525