Systematic Evaluation and Guidelines for Segment Anything Model in Surgical Video Analysis
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914171035058176 |
|---|---|
| author | Yuan, Cheng Jiang, Jian Yang, Kunyi Wu, Lv Wang, Rui Meng, Zi Ping, Haonan Xu, Ziyu Zhou, Yifan Song, Wanli Wang, Hesheng Jin, Yueming Dou, Qi Ban, Yutong |
| author_facet | Yuan, Cheng Jiang, Jian Yang, Kunyi Wu, Lv Wang, Rui Meng, Zi Ping, Haonan Xu, Ziyu Zhou, Yifan Song, Wanli Wang, Hesheng Jin, Yueming Dou, Qi Ban, Yutong |
| contents | Surgical video segmentation is critical for AI to interpret spatial-temporal dynamics in surgery, yet model performance is constrained by limited annotated data. The SAM2 model, pretrained on natural videos, offers potential for zero-shot surgical segmentation, but its applicability in complex surgical environments, with challenges like tissue deformation and instrument variability, remains unexplored. We present the first comprehensive evaluation of the zero-shot capability of SAM2 in 9 surgical datasets (17 surgery types), covering laparoscopic, endoscopic, and robotic procedures. We analyze various prompting (points, boxes, mask) and {finetuning (dense, sparse) strategies}, robustness to surgical challenges, and generalization across procedures and anatomies. Key findings reveal that while SAM2 demonstrates notable zero-shot adaptability in structured scenarios (e.g., instrument segmentation, {multi-organ segmentation}, and scene segmentation), its performance varies under dynamic surgical conditions, highlighting gaps in handling temporal coherence and domain-specific artifacts. These results highlight future pathways to adaptive data-efficient solutions for the surgical data science field. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_00525 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Systematic Evaluation and Guidelines for Segment Anything Model in Surgical Video Analysis Yuan, Cheng Jiang, Jian Yang, Kunyi Wu, Lv Wang, Rui Meng, Zi Ping, Haonan Xu, Ziyu Zhou, Yifan Song, Wanli Wang, Hesheng Jin, Yueming Dou, Qi Ban, Yutong Computer Vision and Pattern Recognition Surgical video segmentation is critical for AI to interpret spatial-temporal dynamics in surgery, yet model performance is constrained by limited annotated data. The SAM2 model, pretrained on natural videos, offers potential for zero-shot surgical segmentation, but its applicability in complex surgical environments, with challenges like tissue deformation and instrument variability, remains unexplored. We present the first comprehensive evaluation of the zero-shot capability of SAM2 in 9 surgical datasets (17 surgery types), covering laparoscopic, endoscopic, and robotic procedures. We analyze various prompting (points, boxes, mask) and {finetuning (dense, sparse) strategies}, robustness to surgical challenges, and generalization across procedures and anatomies. Key findings reveal that while SAM2 demonstrates notable zero-shot adaptability in structured scenarios (e.g., instrument segmentation, {multi-organ segmentation}, and scene segmentation), its performance varies under dynamic surgical conditions, highlighting gaps in handling temporal coherence and domain-specific artifacts. These results highlight future pathways to adaptive data-efficient solutions for the surgical data science field. |
| title | Systematic Evaluation and Guidelines for Segment Anything Model in Surgical Video Analysis |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2501.00525 |