Label-supervised surgical instrument segmentation using temporal equivariance and semantic continuity

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Qiyuan, Liu, Yanzhe, Zhao, Shang, Liu, Rong, Zhou, S. Kevin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911520419479552
author Wang, Qiyuan
Liu, Yanzhe
Zhao, Shang
Liu, Rong
Zhou, S. Kevin
author_facet Wang, Qiyuan
Liu, Yanzhe
Zhao, Shang
Liu, Rong
Zhou, S. Kevin
contents For robotic surgical videos, instrument presence annotations are typically recorded with video streams, which offering the potential to reduce the manually annotated costs for segmentation. However, weakly supervised surgical instrument segmentation with only instrument presence labels has been rarely explored in surgical domain due to the highly under-constrained challenges. Temporal properties can enhance representation learning by capturing sequential dependencies and patterns over time even in incomplete supervision situations. From this, we take the inherent temporal attributes of surgical video into account and extend a two-stage weakly supervised segmentation paradigm from different perspectives. Firstly, we make temporal equivariance constraint to enhance pixel-wise temporal consistency between adjacent features. Secondly, we constrain class-aware semantic continuity between global and local regions across temporal dimension. Finally, we generate temporal-enhanced pseudo masks from consecutive frames to suppress irrelevant regions. Extensive experiments are validated on two surgical video datasets, including one cholecystectomy surgery benchmark and one real robotic left lateral segment liver surgery dataset. We annotate instance-wise instrument labels with fixed time-steps which are double checked by a clinician with 3-years experience to evaluate segmentation results. Experimental results demonstrate the promising performances of our method, which consistently achieves comparable or favorable results with previous state-of-the-art approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2403_09551
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Label-supervised surgical instrument segmentation using temporal equivariance and semantic continuity
Wang, Qiyuan
Liu, Yanzhe
Zhao, Shang
Liu, Rong
Zhou, S. Kevin
Computer Vision and Pattern Recognition
For robotic surgical videos, instrument presence annotations are typically recorded with video streams, which offering the potential to reduce the manually annotated costs for segmentation. However, weakly supervised surgical instrument segmentation with only instrument presence labels has been rarely explored in surgical domain due to the highly under-constrained challenges. Temporal properties can enhance representation learning by capturing sequential dependencies and patterns over time even in incomplete supervision situations. From this, we take the inherent temporal attributes of surgical video into account and extend a two-stage weakly supervised segmentation paradigm from different perspectives. Firstly, we make temporal equivariance constraint to enhance pixel-wise temporal consistency between adjacent features. Secondly, we constrain class-aware semantic continuity between global and local regions across temporal dimension. Finally, we generate temporal-enhanced pseudo masks from consecutive frames to suppress irrelevant regions. Extensive experiments are validated on two surgical video datasets, including one cholecystectomy surgery benchmark and one real robotic left lateral segment liver surgery dataset. We annotate instance-wise instrument labels with fixed time-steps which are double checked by a clinician with 3-years experience to evaluate segmentation results. Experimental results demonstrate the promising performances of our method, which consistently achieves comparable or favorable results with previous state-of-the-art approaches.
title Label-supervised surgical instrument segmentation using temporal equivariance and semantic continuity
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.09551