Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Kun, Srivastav, Vinkle, Yu, Tong, Lavanchy, Joel L., Marescaux, Jacques, Mascagni, Pietro, Navab, Nassir, Padoy, Nicolas
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911007616532480
author Yuan, Kun
Srivastav, Vinkle
Yu, Tong
Lavanchy, Joel L.
Marescaux, Jacques
Mascagni, Pietro
Navab, Nassir
Padoy, Nicolas
author_facet Yuan, Kun
Srivastav, Vinkle
Yu, Tong
Lavanchy, Joel L.
Marescaux, Jacques
Mascagni, Pietro
Navab, Nassir
Padoy, Nicolas
contents Recent advancements in surgical computer vision applications have been driven by vision-only models, which do not explicitly integrate the rich semantics of language into their design. These methods rely on manually annotated surgical videos to predict a fixed set of object categories, limiting their generalizability to unseen surgical procedures and downstream tasks. In this work, we put forward the idea that the surgical video lectures available through open surgical e-learning platforms can provide effective vision and language supervisory signals for multi-modal representation learning without relying on manual annotations. We address the surgery-specific linguistic challenges present in surgical video lectures by employing multiple complementary automatic speech recognition systems to generate text transcriptions. We then present a novel method, SurgVLP - Surgical Vision Language Pre-training, for multi-modal representation learning. Extensive experiments across diverse surgical procedures and tasks demonstrate that the multi-modal representations learned by SurgVLP exhibit strong transferability and adaptability in surgical video analysis. Furthermore, our zero-shot evaluations highlight SurgVLP's potential as a general-purpose foundation model for surgical workflow analysis, reducing the reliance on extensive manual annotations for downstream tasks, and facilitating adaptation methods such as few-shot learning to build a scalable and data-efficient solution for various downstream surgical applications. The [training code](https://github.com/CAMMA-public/PeskaVLP) and [weights](https://github.com/CAMMA-public/SurgVLP) are public.
format Preprint
id arxiv_https___arxiv_org_abs_2307_15220
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
Yuan, Kun
Srivastav, Vinkle
Yu, Tong
Lavanchy, Joel L.
Marescaux, Jacques
Mascagni, Pietro
Navab, Nassir
Padoy, Nicolas
Computer Vision and Pattern Recognition
Artificial Intelligence
Recent advancements in surgical computer vision applications have been driven by vision-only models, which do not explicitly integrate the rich semantics of language into their design. These methods rely on manually annotated surgical videos to predict a fixed set of object categories, limiting their generalizability to unseen surgical procedures and downstream tasks. In this work, we put forward the idea that the surgical video lectures available through open surgical e-learning platforms can provide effective vision and language supervisory signals for multi-modal representation learning without relying on manual annotations. We address the surgery-specific linguistic challenges present in surgical video lectures by employing multiple complementary automatic speech recognition systems to generate text transcriptions. We then present a novel method, SurgVLP - Surgical Vision Language Pre-training, for multi-modal representation learning. Extensive experiments across diverse surgical procedures and tasks demonstrate that the multi-modal representations learned by SurgVLP exhibit strong transferability and adaptability in surgical video analysis. Furthermore, our zero-shot evaluations highlight SurgVLP's potential as a general-purpose foundation model for surgical workflow analysis, reducing the reliance on extensive manual annotations for downstream tasks, and facilitating adaptation methods such as few-shot learning to build a scalable and data-efficient solution for various downstream surgical applications. The [training code](https://github.com/CAMMA-public/PeskaVLP) and [weights](https://github.com/CAMMA-public/SurgVLP) are public.
title Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2307.15220