CataractSAM-2: A Domain-Adapted Model for Anterior Segment Surgery Segmentation and Scalable Ground-Truth Annotation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Eslami, Mohammad, Ganeshkumar, Dhanvinkumar, Kazeminasab, Saber, Morley, Michael G., Boland, Michael V., Lin, Michael M., Miller, John B., Friedman, David S., Zebardast, Nazlee, Sobrin, Lucia, Elze, Tobias
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911537683234816
author Eslami, Mohammad
Ganeshkumar, Dhanvinkumar
Kazeminasab, Saber
Morley, Michael G.
Boland, Michael V.
Lin, Michael M.
Miller, John B.
Friedman, David S.
Zebardast, Nazlee
Sobrin, Lucia
Elze, Tobias
author_facet Eslami, Mohammad
Ganeshkumar, Dhanvinkumar
Kazeminasab, Saber
Morley, Michael G.
Boland, Michael V.
Lin, Michael M.
Miller, John B.
Friedman, David S.
Zebardast, Nazlee
Sobrin, Lucia
Elze, Tobias
contents We present CataractSAM-2, a domain-adapted extension of Meta's Segment Anything Model 2, designed for real-time semantic segmentation of cataract ophthalmic surgery videos with high accuracy. Positioned at the intersection of computer vision and medical robotics, CataractSAM-2 enables precise intraoperative perception crucial for robotic-assisted and computer-guided surgical systems. Furthermore, to alleviate the burden of manual labeling, we introduce an interactive annotation framework that combines sparse prompts with video-based mask propagation. This tool significantly reduces annotation time and facilitates the scalable creation of high-quality ground-truth masks, accelerating dataset development for ocular anterior segment surgeries. We also demonstrate the model's strong zero-shot generalization to glaucoma trabeculectomy procedures, confirming its cross-procedural utility and potential for broader surgical applications. The trained model and annotation toolkit are released as open-source resources, establishing CataractSAM-2 as a foundation for expanding anterior ophthalmic surgical datasets and advancing real-time AI-driven solutions in medical robotics, as well as surgical video understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2603_21566
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CataractSAM-2: A Domain-Adapted Model for Anterior Segment Surgery Segmentation and Scalable Ground-Truth Annotation
Eslami, Mohammad
Ganeshkumar, Dhanvinkumar
Kazeminasab, Saber
Morley, Michael G.
Boland, Michael V.
Lin, Michael M.
Miller, John B.
Friedman, David S.
Zebardast, Nazlee
Sobrin, Lucia
Elze, Tobias
Computer Vision and Pattern Recognition
Artificial Intelligence
Databases
Machine Learning
Robotics
We present CataractSAM-2, a domain-adapted extension of Meta's Segment Anything Model 2, designed for real-time semantic segmentation of cataract ophthalmic surgery videos with high accuracy. Positioned at the intersection of computer vision and medical robotics, CataractSAM-2 enables precise intraoperative perception crucial for robotic-assisted and computer-guided surgical systems. Furthermore, to alleviate the burden of manual labeling, we introduce an interactive annotation framework that combines sparse prompts with video-based mask propagation. This tool significantly reduces annotation time and facilitates the scalable creation of high-quality ground-truth masks, accelerating dataset development for ocular anterior segment surgeries. We also demonstrate the model's strong zero-shot generalization to glaucoma trabeculectomy procedures, confirming its cross-procedural utility and potential for broader surgical applications. The trained model and annotation toolkit are released as open-source resources, establishing CataractSAM-2 as a foundation for expanding anterior ophthalmic surgical datasets and advancing real-time AI-driven solutions in medical robotics, as well as surgical video understanding.
title CataractSAM-2: A Domain-Adapted Model for Anterior Segment Surgery Segmentation and Scalable Ground-Truth Annotation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Databases
Machine Learning
Robotics
url https://arxiv.org/abs/2603.21566