Transferring Relative Monocular Depth to Surgical Vision with Temporal Consistency

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Budd, Charlie, Vercauteren, Tom
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929550658633728
author Budd, Charlie
Vercauteren, Tom
author_facet Budd, Charlie
Vercauteren, Tom
contents Relative monocular depth, inferring depth up to shift and scale from a single image, is an active research topic. Recent deep learning models, trained on large and varied meta-datasets, now provide excellent performance in the domain of natural images. However, few datasets exist which provide ground truth depth for endoscopic images, making training such models from scratch unfeasible. This work investigates the transfer of these models into the surgical domain, and presents an effective and simple way to improve on standard supervision through the use of temporal consistency self-supervision. We show temporal consistency significantly improves supervised training alone when transferring to the low-data regime of endoscopy, and outperforms the prevalent self-supervision technique for this task. In addition we show our method drastically outperforms the state-of-the-art method from within the domain of endoscopy. We also release our code, model and ensembled meta-dataset, Meta-MED, establishing a strong benchmark for future work.
format Preprint
id arxiv_https___arxiv_org_abs_2403_06683
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Transferring Relative Monocular Depth to Surgical Vision with Temporal Consistency
Budd, Charlie
Vercauteren, Tom
Computer Vision and Pattern Recognition
Relative monocular depth, inferring depth up to shift and scale from a single image, is an active research topic. Recent deep learning models, trained on large and varied meta-datasets, now provide excellent performance in the domain of natural images. However, few datasets exist which provide ground truth depth for endoscopic images, making training such models from scratch unfeasible. This work investigates the transfer of these models into the surgical domain, and presents an effective and simple way to improve on standard supervision through the use of temporal consistency self-supervision. We show temporal consistency significantly improves supervised training alone when transferring to the low-data regime of endoscopy, and outperforms the prevalent self-supervision technique for this task. In addition we show our method drastically outperforms the state-of-the-art method from within the domain of endoscopy. We also release our code, model and ensembled meta-dataset, Meta-MED, establishing a strong benchmark for future work.
title Transferring Relative Monocular Depth to Surgical Vision with Temporal Consistency
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.06683