SemiVT-Surge: Semi-Supervised Video Transformer for Surgical Phase Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yiping, de Jong, Ronald, Nasirihaghighi, Sahar, Jaspers, Tim, van Jaarsveld, Romy, Kuiper, Gino, van Hillegersberg, Richard, van der Sommen, Fons, Ruurda, Jelle, Breeuwer, Marcel, Khalil, Yasmina Al
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915318174056448
author Li, Yiping
de Jong, Ronald
Nasirihaghighi, Sahar
Jaspers, Tim
van Jaarsveld, Romy
Kuiper, Gino
van Hillegersberg, Richard
van der Sommen, Fons
Ruurda, Jelle
Breeuwer, Marcel
Khalil, Yasmina Al
author_facet Li, Yiping
de Jong, Ronald
Nasirihaghighi, Sahar
Jaspers, Tim
van Jaarsveld, Romy
Kuiper, Gino
van Hillegersberg, Richard
van der Sommen, Fons
Ruurda, Jelle
Breeuwer, Marcel
Khalil, Yasmina Al
contents Accurate surgical phase recognition is crucial for computer-assisted interventions and surgical video analysis. Annotating long surgical videos is labor-intensive, driving research toward leveraging unlabeled data for strong performance with minimal annotations. Although self-supervised learning has gained popularity by enabling large-scale pretraining followed by fine-tuning on small labeled subsets, semi-supervised approaches remain largely underexplored in the surgical domain. In this work, we propose a video transformer-based model with a robust pseudo-labeling framework. Our method incorporates temporal consistency regularization for unlabeled data and contrastive learning with class prototypes, which leverages both labeled data and pseudo-labels to refine the feature space. Through extensive experiments on the private RAMIE (Robot-Assisted Minimally Invasive Esophagectomy) dataset and the public Cholec80 dataset, we demonstrate the effectiveness of our approach. By incorporating unlabeled data, we achieve state-of-the-art performance on RAMIE with a 4.9% accuracy increase and obtain comparable results to full supervision while using only 1/4 of the labeled data on Cholec80. Our findings establish a strong benchmark for semi-supervised surgical phase recognition, paving the way for future research in this domain.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01471
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SemiVT-Surge: Semi-Supervised Video Transformer for Surgical Phase Recognition
Li, Yiping
de Jong, Ronald
Nasirihaghighi, Sahar
Jaspers, Tim
van Jaarsveld, Romy
Kuiper, Gino
van Hillegersberg, Richard
van der Sommen, Fons
Ruurda, Jelle
Breeuwer, Marcel
Khalil, Yasmina Al
Computer Vision and Pattern Recognition
Accurate surgical phase recognition is crucial for computer-assisted interventions and surgical video analysis. Annotating long surgical videos is labor-intensive, driving research toward leveraging unlabeled data for strong performance with minimal annotations. Although self-supervised learning has gained popularity by enabling large-scale pretraining followed by fine-tuning on small labeled subsets, semi-supervised approaches remain largely underexplored in the surgical domain. In this work, we propose a video transformer-based model with a robust pseudo-labeling framework. Our method incorporates temporal consistency regularization for unlabeled data and contrastive learning with class prototypes, which leverages both labeled data and pseudo-labels to refine the feature space. Through extensive experiments on the private RAMIE (Robot-Assisted Minimally Invasive Esophagectomy) dataset and the public Cholec80 dataset, we demonstrate the effectiveness of our approach. By incorporating unlabeled data, we achieve state-of-the-art performance on RAMIE with a 4.9% accuracy increase and obtain comparable results to full supervision while using only 1/4 of the labeled data on Cholec80. Our findings establish a strong benchmark for semi-supervised surgical phase recognition, paving the way for future research in this domain.
title SemiVT-Surge: Semi-Supervised Video Transformer for Surgical Phase Recognition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.01471