Leveraging Self-Supervised Learning for Fetal Cardiac Planes Classification using Ultrasound Scan Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Benjamin, Joseph Geo, Asokan, Mothilal, Alhosani, Amna, Alasmawi, Hussain, Diehl, Werner Gerhard, Bricker, Leanne, Nandakumar, Karthik, Yaqub, Mohammad
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916341655535616
author Benjamin, Joseph Geo
Asokan, Mothilal
Alhosani, Amna
Alasmawi, Hussain
Diehl, Werner Gerhard
Bricker, Leanne
Nandakumar, Karthik
Yaqub, Mohammad
author_facet Benjamin, Joseph Geo
Asokan, Mothilal
Alhosani, Amna
Alasmawi, Hussain
Diehl, Werner Gerhard
Bricker, Leanne
Nandakumar, Karthik
Yaqub, Mohammad
contents Self-supervised learning (SSL) methods are popular since they can address situations with limited annotated data by directly utilising the underlying data distribution. However, the adoption of such methods is not explored enough in ultrasound (US) imaging, especially for fetal assessment. We investigate the potential of dual-encoder SSL in utilizing unlabelled US video data to improve the performance of challenging downstream Standard Fetal Cardiac Planes (SFCP) classification using limited labelled 2D US images. We study 7 SSL approaches based on reconstruction, contrastive loss, distillation, and information theory and evaluate them extensively on a large private US dataset. Our observations and findings are consolidated from more than 500 downstream training experiments under different settings. Our primary observation shows that for SSL training, the variance of the dataset is more crucial than its size because it allows the model to learn generalisable representations, which improve the performance of downstream tasks. Overall, the BarlowTwins method shows robust performance, irrespective of the training settings and data variations, when used as an initialisation for downstream tasks. Notably, full fine-tuning with 1% of labelled data outperforms ImageNet initialisation by 12% in F1-score and outperforms other SSL initialisations by at least 4% in F1-score, thus making it a promising candidate for transfer learning from US video to image data.
format Preprint
id arxiv_https___arxiv_org_abs_2407_21738
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging Self-Supervised Learning for Fetal Cardiac Planes Classification using Ultrasound Scan Videos
Benjamin, Joseph Geo
Asokan, Mothilal
Alhosani, Amna
Alasmawi, Hussain
Diehl, Werner Gerhard
Bricker, Leanne
Nandakumar, Karthik
Yaqub, Mohammad
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Self-supervised learning (SSL) methods are popular since they can address situations with limited annotated data by directly utilising the underlying data distribution. However, the adoption of such methods is not explored enough in ultrasound (US) imaging, especially for fetal assessment. We investigate the potential of dual-encoder SSL in utilizing unlabelled US video data to improve the performance of challenging downstream Standard Fetal Cardiac Planes (SFCP) classification using limited labelled 2D US images. We study 7 SSL approaches based on reconstruction, contrastive loss, distillation, and information theory and evaluate them extensively on a large private US dataset. Our observations and findings are consolidated from more than 500 downstream training experiments under different settings. Our primary observation shows that for SSL training, the variance of the dataset is more crucial than its size because it allows the model to learn generalisable representations, which improve the performance of downstream tasks. Overall, the BarlowTwins method shows robust performance, irrespective of the training settings and data variations, when used as an initialisation for downstream tasks. Notably, full fine-tuning with 1% of labelled data outperforms ImageNet initialisation by 12% in F1-score and outperforms other SSL initialisations by at least 4% in F1-score, thus making it a promising candidate for transfer learning from US video to image data.
title Leveraging Self-Supervised Learning for Fetal Cardiac Planes Classification using Ultrasound Scan Videos
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2407.21738