A generalizable foundation model for intraoperative understanding across surgical procedures

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Park, Kanggil, Jeon, Yongjun, Lim, Soyoung, Park, Seonmin, Shin, Jongmin, Kim, Jung Yong, An, Sehyeon, Rhu, Jinsoo, Kim, Jongman, Choi, Gyu-Seong, Oh, Namkee, Jung, Kyu-Hwan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914329681461248
author Park, Kanggil
Jeon, Yongjun
Lim, Soyoung
Park, Seonmin
Shin, Jongmin
Kim, Jung Yong
An, Sehyeon
Rhu, Jinsoo
Kim, Jongman
Choi, Gyu-Seong
Oh, Namkee
Jung, Kyu-Hwan
author_facet Park, Kanggil
Jeon, Yongjun
Lim, Soyoung
Park, Seonmin
Shin, Jongmin
Kim, Jung Yong
An, Sehyeon
Rhu, Jinsoo
Kim, Jongman
Choi, Gyu-Seong
Oh, Namkee
Jung, Kyu-Hwan
contents In minimally invasive surgery, clinical decisions depend on real-time visual interpretation, yet intraoperative perception varies substantially across surgeons and procedures. This variability limits consistent assessment, training, and the development of reliable artificial intelligence systems, as most surgical AI models are designed for narrowly defined tasks and do not generalize across procedures or institutions. Here we introduce ZEN, a generalizable foundation model for intraoperative surgical video understanding trained on more than 4 million frames from over 21 procedures using a self-supervised multi-teacher distillation framework. We curated a large and diverse dataset and systematically evaluated multiple representation learning strategies within a unified benchmark. Across 20 downstream tasks and full fine-tuning, frozen-backbone, few-shot and zero-shot settings, ZEN consistently outperforms existing surgical foundation models and demonstrates robust cross-procedure generalization. These results suggest a step toward unified representations for surgical scene understanding and support future applications in intraoperative assistance and surgical training assessment.
format Preprint
id arxiv_https___arxiv_org_abs_2602_13633
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A generalizable foundation model for intraoperative understanding across surgical procedures
Park, Kanggil
Jeon, Yongjun
Lim, Soyoung
Park, Seonmin
Shin, Jongmin
Kim, Jung Yong
An, Sehyeon
Rhu, Jinsoo
Kim, Jongman
Choi, Gyu-Seong
Oh, Namkee
Jung, Kyu-Hwan
Computer Vision and Pattern Recognition
In minimally invasive surgery, clinical decisions depend on real-time visual interpretation, yet intraoperative perception varies substantially across surgeons and procedures. This variability limits consistent assessment, training, and the development of reliable artificial intelligence systems, as most surgical AI models are designed for narrowly defined tasks and do not generalize across procedures or institutions. Here we introduce ZEN, a generalizable foundation model for intraoperative surgical video understanding trained on more than 4 million frames from over 21 procedures using a self-supervised multi-teacher distillation framework. We curated a large and diverse dataset and systematically evaluated multiple representation learning strategies within a unified benchmark. Across 20 downstream tasks and full fine-tuning, frozen-backbone, few-shot and zero-shot settings, ZEN consistently outperforms existing surgical foundation models and demonstrates robust cross-procedure generalization. These results suggest a step toward unified representations for surgical scene understanding and support future applications in intraoperative assistance and surgical training assessment.
title A generalizable foundation model for intraoperative understanding across surgical procedures
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.13633