General surgery vision transformer: A video pre-trained foundation model for general surgery

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schmidgall, Samuel, Kim, Ji Woong, Jopling, Jeffrey, Krieger, Axel
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916204201902080
author Schmidgall, Samuel
Kim, Ji Woong
Jopling, Jeffrey
Krieger, Axel
author_facet Schmidgall, Samuel
Kim, Ji Woong
Jopling, Jeffrey
Krieger, Axel
contents The absence of openly accessible data and specialized foundation models is a major barrier for computational research in surgery. Toward this, (i) we open-source the largest dataset of general surgery videos to-date, consisting of 680 hours of surgical videos, including data from robotic and laparoscopic techniques across 28 procedures; (ii) we propose a technique for video pre-training a general surgery vision transformer (GSViT) on surgical videos based on forward video prediction that can run in real-time for surgical applications, toward which we open-source the code and weights of GSViT; (iii) we also release code and weights for procedure-specific fine-tuned versions of GSViT across 10 procedures; (iv) we demonstrate the performance of GSViT on the Cholec80 phase annotation task, displaying improved performance over state-of-the-art single frame predictors.
format Preprint
id arxiv_https___arxiv_org_abs_2403_05949
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle General surgery vision transformer: A video pre-trained foundation model for general surgery
Schmidgall, Samuel
Kim, Ji Woong
Jopling, Jeffrey
Krieger, Axel
Computer Vision and Pattern Recognition
Machine Learning
Tissues and Organs
The absence of openly accessible data and specialized foundation models is a major barrier for computational research in surgery. Toward this, (i) we open-source the largest dataset of general surgery videos to-date, consisting of 680 hours of surgical videos, including data from robotic and laparoscopic techniques across 28 procedures; (ii) we propose a technique for video pre-training a general surgery vision transformer (GSViT) on surgical videos based on forward video prediction that can run in real-time for surgical applications, toward which we open-source the code and weights of GSViT; (iii) we also release code and weights for procedure-specific fine-tuned versions of GSViT across 10 procedures; (iv) we demonstrate the performance of GSViT on the Cholec80 phase annotation task, displaying improved performance over state-of-the-art single frame predictors.
title General surgery vision transformer: A video pre-trained foundation model for general surgery
topic Computer Vision and Pattern Recognition
Machine Learning
Tissues and Organs
url https://arxiv.org/abs/2403.05949