Systematic Evaluation and Guidelines for Segment Anything Model in Surgical Video Analysis

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yuan, Cheng, Jiang, Jian, Yang, Kunyi, Wu, Lv, Wang, Rui, Meng, Zi, Ping, Haonan, Xu, Ziyu, Zhou, Yifan, Song, Wanli, Wang, Hesheng, Jin, Yueming, Dou, Qi, Ban, Yutong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914171035058176
author Yuan, Cheng
Jiang, Jian
Yang, Kunyi
Wu, Lv
Wang, Rui
Meng, Zi
Ping, Haonan
Xu, Ziyu
Zhou, Yifan
Song, Wanli
Wang, Hesheng
Jin, Yueming
Dou, Qi
Ban, Yutong
author_facet Yuan, Cheng
Jiang, Jian
Yang, Kunyi
Wu, Lv
Wang, Rui
Meng, Zi
Ping, Haonan
Xu, Ziyu
Zhou, Yifan
Song, Wanli
Wang, Hesheng
Jin, Yueming
Dou, Qi
Ban, Yutong
contents Surgical video segmentation is critical for AI to interpret spatial-temporal dynamics in surgery, yet model performance is constrained by limited annotated data. The SAM2 model, pretrained on natural videos, offers potential for zero-shot surgical segmentation, but its applicability in complex surgical environments, with challenges like tissue deformation and instrument variability, remains unexplored. We present the first comprehensive evaluation of the zero-shot capability of SAM2 in 9 surgical datasets (17 surgery types), covering laparoscopic, endoscopic, and robotic procedures. We analyze various prompting (points, boxes, mask) and {finetuning (dense, sparse) strategies}, robustness to surgical challenges, and generalization across procedures and anatomies. Key findings reveal that while SAM2 demonstrates notable zero-shot adaptability in structured scenarios (e.g., instrument segmentation, {multi-organ segmentation}, and scene segmentation), its performance varies under dynamic surgical conditions, highlighting gaps in handling temporal coherence and domain-specific artifacts. These results highlight future pathways to adaptive data-efficient solutions for the surgical data science field.
format Preprint
id arxiv_https___arxiv_org_abs_2501_00525
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Systematic Evaluation and Guidelines for Segment Anything Model in Surgical Video Analysis
Yuan, Cheng
Jiang, Jian
Yang, Kunyi
Wu, Lv
Wang, Rui
Meng, Zi
Ping, Haonan
Xu, Ziyu
Zhou, Yifan
Song, Wanli
Wang, Hesheng
Jin, Yueming
Dou, Qi
Ban, Yutong
Computer Vision and Pattern Recognition
Surgical video segmentation is critical for AI to interpret spatial-temporal dynamics in surgery, yet model performance is constrained by limited annotated data. The SAM2 model, pretrained on natural videos, offers potential for zero-shot surgical segmentation, but its applicability in complex surgical environments, with challenges like tissue deformation and instrument variability, remains unexplored. We present the first comprehensive evaluation of the zero-shot capability of SAM2 in 9 surgical datasets (17 surgery types), covering laparoscopic, endoscopic, and robotic procedures. We analyze various prompting (points, boxes, mask) and {finetuning (dense, sparse) strategies}, robustness to surgical challenges, and generalization across procedures and anatomies. Key findings reveal that while SAM2 demonstrates notable zero-shot adaptability in structured scenarios (e.g., instrument segmentation, {multi-organ segmentation}, and scene segmentation), its performance varies under dynamic surgical conditions, highlighting gaps in handling temporal coherence and domain-specific artifacts. These results highlight future pathways to adaptive data-efficient solutions for the surgical data science field.
title Systematic Evaluation and Guidelines for Segment Anything Model in Surgical Video Analysis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.00525