Self-Supervised Ultrasound-Video Segmentation with Feature Prediction and 3D Localised Loss

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ellis, Edward, Mendel, Robert, Bulpitt, Andrew, Parsa, Nasim, Byrne, Michael F, Ali, Sharib
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908464683417600
author Ellis, Edward
Mendel, Robert
Bulpitt, Andrew
Parsa, Nasim
Byrne, Michael F
Ali, Sharib
author_facet Ellis, Edward
Mendel, Robert
Bulpitt, Andrew
Parsa, Nasim
Byrne, Michael F
Ali, Sharib
contents Acquiring and annotating large datasets in ultrasound imaging is challenging due to low contrast, high noise, and susceptibility to artefacts. This process requires significant time and clinical expertise. Self-supervised learning (SSL) offers a promising solution by leveraging unlabelled data to learn useful representations, enabling improved segmentation performance when annotated data is limited. Recent state-of-the-art developments in SSL for video data include V-JEPA, a framework solely based on feature prediction, avoiding pixel level reconstruction or negative samples. We hypothesise that V-JEPA is well-suited to ultrasound imaging, as it is less sensitive to noisy pixel-level detail while effectively leveraging temporal information. To the best of our knowledge, this is the first study to adopt V-JEPA for ultrasound video data. Similar to other patch-based masking SSL techniques such as VideoMAE, V-JEPA is well-suited to ViT-based models. However, ViTs can underperform on small medical datasets due to lack of inductive biases, limited spatial locality and absence of hierarchical feature learning. To improve locality understanding, we propose a novel 3D localisation auxiliary task to improve locality in ViT representations during V-JEPA pre-training. Our results show V-JEPA with our auxiliary task improves segmentation performance significantly across various frozen encoder configurations, with gains up to 3.4\% using 100\% and up to 8.35\% using only 10\% of the training data.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18424
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-Supervised Ultrasound-Video Segmentation with Feature Prediction and 3D Localised Loss
Ellis, Edward
Mendel, Robert
Bulpitt, Andrew
Parsa, Nasim
Byrne, Michael F
Ali, Sharib
Computer Vision and Pattern Recognition
Acquiring and annotating large datasets in ultrasound imaging is challenging due to low contrast, high noise, and susceptibility to artefacts. This process requires significant time and clinical expertise. Self-supervised learning (SSL) offers a promising solution by leveraging unlabelled data to learn useful representations, enabling improved segmentation performance when annotated data is limited. Recent state-of-the-art developments in SSL for video data include V-JEPA, a framework solely based on feature prediction, avoiding pixel level reconstruction or negative samples. We hypothesise that V-JEPA is well-suited to ultrasound imaging, as it is less sensitive to noisy pixel-level detail while effectively leveraging temporal information. To the best of our knowledge, this is the first study to adopt V-JEPA for ultrasound video data. Similar to other patch-based masking SSL techniques such as VideoMAE, V-JEPA is well-suited to ViT-based models. However, ViTs can underperform on small medical datasets due to lack of inductive biases, limited spatial locality and absence of hierarchical feature learning. To improve locality understanding, we propose a novel 3D localisation auxiliary task to improve locality in ViT representations during V-JEPA pre-training. Our results show V-JEPA with our auxiliary task improves segmentation performance significantly across various frozen encoder configurations, with gains up to 3.4\% using 100\% and up to 8.35\% using only 10\% of the training data.
title Self-Supervised Ultrasound-Video Segmentation with Feature Prediction and 3D Localised Loss
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.18424