Surgical Foundation Model Leveraging Compression and Entropy Maximization for Image-Guided Surgical Assistance

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yin, Lianhao, Meireles, Ozanan, Rosman, Guy, Rus, Daniela
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913871870033920
author Yin, Lianhao
Meireles, Ozanan
Rosman, Guy
Rus, Daniela
author_facet Yin, Lianhao
Meireles, Ozanan
Rosman, Guy
Rus, Daniela
contents Real-time video understanding is critical to guide procedures in minimally invasive surgery (MIS). However, supervised learning approaches require large, annotated datasets that are scarce due to annotation efforts that are prohibitive, e.g., in medical fields. Although self-supervision methods can address such limitations, current self-supervised methods often fail to capture structural and physical information in a form that generalizes across tasks. We propose Compress-to-Explore (C2E), a novel self-supervised framework that leverages Kolmogorov complexity to learn compact, informative representations from surgical videos. C2E uses entropy-maximizing decoders to compress images while preserving clinically relevant details, improving encoder performance without labeled data. Trained on large-scale unlabeled surgical datasets, C2E demonstrates strong generalization across a variety of surgical ML tasks, such as workflow classification, tool-tissue interaction classification, segmentation, and diagnosis tasks, providing improved performance as a surgical visual foundation model. As we further show in the paper, the model's internal compact representation better disentangles features from different structural parts of images. The resulting performance improvements highlight the yet untapped potential of self-supervised learning to enhance surgical AI and improve outcomes in MIS.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01980
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Surgical Foundation Model Leveraging Compression and Entropy Maximization for Image-Guided Surgical Assistance
Yin, Lianhao
Meireles, Ozanan
Rosman, Guy
Rus, Daniela
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Real-time video understanding is critical to guide procedures in minimally invasive surgery (MIS). However, supervised learning approaches require large, annotated datasets that are scarce due to annotation efforts that are prohibitive, e.g., in medical fields. Although self-supervision methods can address such limitations, current self-supervised methods often fail to capture structural and physical information in a form that generalizes across tasks. We propose Compress-to-Explore (C2E), a novel self-supervised framework that leverages Kolmogorov complexity to learn compact, informative representations from surgical videos. C2E uses entropy-maximizing decoders to compress images while preserving clinically relevant details, improving encoder performance without labeled data. Trained on large-scale unlabeled surgical datasets, C2E demonstrates strong generalization across a variety of surgical ML tasks, such as workflow classification, tool-tissue interaction classification, segmentation, and diagnosis tasks, providing improved performance as a surgical visual foundation model. As we further show in the paper, the model's internal compact representation better disentangles features from different structural parts of images. The resulting performance improvements highlight the yet untapped potential of self-supervised learning to enhance surgical AI and improve outcomes in MIS.
title Surgical Foundation Model Leveraging Compression and Entropy Maximization for Image-Guided Surgical Assistance
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.01980