Exploring Pre-training Across Domains for Few-Shot Surgical Skill Assessment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Anastasiou, Dimitrios, Caramalau, Razvan, Sirajudeen, Nazir, Boal, Matthew, Edwards, Philip, Collins, Justin, Kelly, John, Sridhar, Ashwin, Tran, Maxine, Mumtaz, Faiz, Pavithran, Nevil, Francis, Nader, Stoyanov, Danail, Mazomenos, Evangelos B.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918139352055808
author Anastasiou, Dimitrios
Caramalau, Razvan
Sirajudeen, Nazir
Boal, Matthew
Edwards, Philip
Collins, Justin
Kelly, John
Sridhar, Ashwin
Tran, Maxine
Mumtaz, Faiz
Pavithran, Nevil
Francis, Nader
Stoyanov, Danail
Mazomenos, Evangelos B.
author_facet Anastasiou, Dimitrios
Caramalau, Razvan
Sirajudeen, Nazir
Boal, Matthew
Edwards, Philip
Collins, Justin
Kelly, John
Sridhar, Ashwin
Tran, Maxine
Mumtaz, Faiz
Pavithran, Nevil
Francis, Nader
Stoyanov, Danail
Mazomenos, Evangelos B.
contents Automated surgical skill assessment (SSA) is a central task in surgical computer vision. Developing robust SSA models is challenging due to the scarcity of skill annotations, which are time-consuming to produce and require expert consensus. Few-shot learning (FSL) offers a scalable alternative enabling model development with minimal supervision, though its success critically depends on effective pre-training. While widely studied for several surgical downstream tasks, pre-training has remained largely unexplored in SSA. In this work, we formulate SSA as a few-shot task and investigate how self-supervised pre-training strategies affect downstream few-shot SSA performance. We annotate a publicly available robotic surgery dataset with Objective Structured Assessment of Technical Skill (OSATS) scores, and evaluate various pre-training sources across three few-shot settings. We quantify domain similarity and analyze how domain gap and the inclusion of procedure-specific data into pre-training influence transferability. Our results show that small but domain-relevant datasets can outperform large scale, less aligned ones, achieving accuracies of 60.16%, 66.03%, and 73.65% in the 1-, 2-, and 5-shot settings, respectively. Moreover, incorporating procedure-specific data into pre-training with a domain-relevant external dataset significantly boosts downstream performance, with an average gain of +1.22% in accuracy and +2.28% in F1-score; however, applying the same strategy with less similar but large-scale sources can instead lead to performance degradation. Code and models are available at https://github.com/anastadimi/ssa-fsl.
format Preprint
id arxiv_https___arxiv_org_abs_2509_09327
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Pre-training Across Domains for Few-Shot Surgical Skill Assessment
Anastasiou, Dimitrios
Caramalau, Razvan
Sirajudeen, Nazir
Boal, Matthew
Edwards, Philip
Collins, Justin
Kelly, John
Sridhar, Ashwin
Tran, Maxine
Mumtaz, Faiz
Pavithran, Nevil
Francis, Nader
Stoyanov, Danail
Mazomenos, Evangelos B.
Computer Vision and Pattern Recognition
Machine Learning
Automated surgical skill assessment (SSA) is a central task in surgical computer vision. Developing robust SSA models is challenging due to the scarcity of skill annotations, which are time-consuming to produce and require expert consensus. Few-shot learning (FSL) offers a scalable alternative enabling model development with minimal supervision, though its success critically depends on effective pre-training. While widely studied for several surgical downstream tasks, pre-training has remained largely unexplored in SSA. In this work, we formulate SSA as a few-shot task and investigate how self-supervised pre-training strategies affect downstream few-shot SSA performance. We annotate a publicly available robotic surgery dataset with Objective Structured Assessment of Technical Skill (OSATS) scores, and evaluate various pre-training sources across three few-shot settings. We quantify domain similarity and analyze how domain gap and the inclusion of procedure-specific data into pre-training influence transferability. Our results show that small but domain-relevant datasets can outperform large scale, less aligned ones, achieving accuracies of 60.16%, 66.03%, and 73.65% in the 1-, 2-, and 5-shot settings, respectively. Moreover, incorporating procedure-specific data into pre-training with a domain-relevant external dataset significantly boosts downstream performance, with an average gain of +1.22% in accuracy and +2.28% in F1-score; however, applying the same strategy with less similar but large-scale sources can instead lead to performance degradation. Code and models are available at https://github.com/anastadimi/ssa-fsl.
title Exploring Pre-training Across Domains for Few-Shot Surgical Skill Assessment
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2509.09327