Adapting Foundation Models for Annotation-Efficient Adnexal Mass Segmentation in Cine Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fati, Francesca, Rota, Alberto, Gregory, Adriana V., Catozzo, Anna, Giuliano, Maria C., Dhar, Mrinal, De Vitis, Luigi, Packard, Annie T., Multinu, Francesco, De Momi, Elena, Langstraat, Carrie L., Kline, Timothy L.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918436650614784
author Fati, Francesca
Rota, Alberto
Gregory, Adriana V.
Catozzo, Anna
Giuliano, Maria C.
Dhar, Mrinal
De Vitis, Luigi
Packard, Annie T.
Multinu, Francesco
De Momi, Elena
Langstraat, Carrie L.
Kline, Timothy L.
author_facet Fati, Francesca
Rota, Alberto
Gregory, Adriana V.
Catozzo, Anna
Giuliano, Maria C.
Dhar, Mrinal
De Vitis, Luigi
Packard, Annie T.
Multinu, Francesco
De Momi, Elena
Langstraat, Carrie L.
Kline, Timothy L.
contents Adnexal mass evaluation via ultrasound is a challenging clinical task, often hindered by subjective interpretation and significant inter-observer variability. While automated segmentation is a foundational step for quantitative risk assessment, traditional fully supervised convolutional architectures frequently require large amounts of pixel-level annotations and struggle with domain shifts common in medical imaging. In this work, we propose a label-efficient segmentation framework that leverages the robust semantic priors of a pretrained DINOv3 foundational vision transformer backbone. By integrating this backbone with a Dense Prediction Transformer (DPT)-style decoder, our model hierarchically reassembles multi-scale features to combine global semantic representations with fine-grained spatial details. Evaluated on a clinical dataset of 7,777 annotated frames from 112 patients, our method achieves state-of-the-art performance compared to established fully supervised baselines, including U-Net, U-Net++, DeepLabV3, and MAnet. Specifically, we obtain a Dice score of 0.945 and improved boundary adherence, reducing the 95th-percentile Hausdorff Distance by 11.4% relative to the strongest convolutional baseline. Furthermore, we conduct an extensive efficiency analysis demonstrating that our DINOv3-based approach retains significantly higher performance under data starvation regimes, maintaining strong results even when trained on only 25% of the data. These results suggest that leveraging large-scale self-supervised foundations provides a promising and data-efficient solution for medical image segmentation in data-constrained clinical environments. Project Repository: https://github.com/FrancescaFati/MESA
format Preprint
id arxiv_https___arxiv_org_abs_2604_08045
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Adapting Foundation Models for Annotation-Efficient Adnexal Mass Segmentation in Cine Images
Fati, Francesca
Rota, Alberto
Gregory, Adriana V.
Catozzo, Anna
Giuliano, Maria C.
Dhar, Mrinal
De Vitis, Luigi
Packard, Annie T.
Multinu, Francesco
De Momi, Elena
Langstraat, Carrie L.
Kline, Timothy L.
Computer Vision and Pattern Recognition
Adnexal mass evaluation via ultrasound is a challenging clinical task, often hindered by subjective interpretation and significant inter-observer variability. While automated segmentation is a foundational step for quantitative risk assessment, traditional fully supervised convolutional architectures frequently require large amounts of pixel-level annotations and struggle with domain shifts common in medical imaging. In this work, we propose a label-efficient segmentation framework that leverages the robust semantic priors of a pretrained DINOv3 foundational vision transformer backbone. By integrating this backbone with a Dense Prediction Transformer (DPT)-style decoder, our model hierarchically reassembles multi-scale features to combine global semantic representations with fine-grained spatial details. Evaluated on a clinical dataset of 7,777 annotated frames from 112 patients, our method achieves state-of-the-art performance compared to established fully supervised baselines, including U-Net, U-Net++, DeepLabV3, and MAnet. Specifically, we obtain a Dice score of 0.945 and improved boundary adherence, reducing the 95th-percentile Hausdorff Distance by 11.4% relative to the strongest convolutional baseline. Furthermore, we conduct an extensive efficiency analysis demonstrating that our DINOv3-based approach retains significantly higher performance under data starvation regimes, maintaining strong results even when trained on only 25% of the data. These results suggest that leveraging large-scale self-supervised foundations provides a promising and data-efficient solution for medical image segmentation in data-constrained clinical environments. Project Repository: https://github.com/FrancescaFati/MESA
title Adapting Foundation Models for Annotation-Efficient Adnexal Mass Segmentation in Cine Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.08045