Self-supervised Learning of Echocardiographic Video Representations via Online Cluster Distillation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mishra, Divyanshu, Salehi, Mohammadreza, Saha, Pramit, Patey, Olga, Papageorghiou, Aris T., Asano, Yuki M., Noble, J. Alison
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914274141536256
author Mishra, Divyanshu
Salehi, Mohammadreza
Saha, Pramit
Patey, Olga
Papageorghiou, Aris T.
Asano, Yuki M.
Noble, J. Alison
author_facet Mishra, Divyanshu
Salehi, Mohammadreza
Saha, Pramit
Patey, Olga
Papageorghiou, Aris T.
Asano, Yuki M.
Noble, J. Alison
contents Self-supervised learning (SSL) has achieved major advances in natural images and video understanding, but challenges remain in domains like echocardiography (heart ultrasound) due to subtle anatomical structures, complex temporal dynamics, and the current lack of domain-specific pre-trained models. Existing SSL approaches such as contrastive, masked modeling, and clustering-based methods struggle with high intersample similarity, sensitivity to low PSNR inputs common in ultrasound, or aggressive augmentations that distort clinically relevant features. We present DISCOVR (Distilled Image Supervision for Cross Modal Video Representation), a self-supervised dual branch framework for cardiac ultrasound video representation learning. DISCOVR combines a clustering-based video encoder that models temporal dynamics with an online image encoder that extracts fine-grained spatial semantics. These branches are connected through a semantic cluster distillation loss that transfers anatomical knowledge from the evolving image encoder to the video encoder, enabling temporally coherent representations enriched with fine-grained semantic understanding.Evaluated on six echocardiography datasets spanning fetal, pediatric, and adult populations, DISCOVR outperforms both specialized video anomaly detection methods and state-of-the-art video-SSL baselines in zero-shot and linear probing setups,achieving superior segmentation transfer and strong downstream performance on clinically relevant tasks such as LVEF prediction. Code available at: https://github.com/mdivyanshu97/DISCOVR
format Preprint
id arxiv_https___arxiv_org_abs_2506_11777
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-supervised Learning of Echocardiographic Video Representations via Online Cluster Distillation
Mishra, Divyanshu
Salehi, Mohammadreza
Saha, Pramit
Patey, Olga
Papageorghiou, Aris T.
Asano, Yuki M.
Noble, J. Alison
Computer Vision and Pattern Recognition
Artificial Intelligence
Computers and Society
Machine Learning
Self-supervised learning (SSL) has achieved major advances in natural images and video understanding, but challenges remain in domains like echocardiography (heart ultrasound) due to subtle anatomical structures, complex temporal dynamics, and the current lack of domain-specific pre-trained models. Existing SSL approaches such as contrastive, masked modeling, and clustering-based methods struggle with high intersample similarity, sensitivity to low PSNR inputs common in ultrasound, or aggressive augmentations that distort clinically relevant features. We present DISCOVR (Distilled Image Supervision for Cross Modal Video Representation), a self-supervised dual branch framework for cardiac ultrasound video representation learning. DISCOVR combines a clustering-based video encoder that models temporal dynamics with an online image encoder that extracts fine-grained spatial semantics. These branches are connected through a semantic cluster distillation loss that transfers anatomical knowledge from the evolving image encoder to the video encoder, enabling temporally coherent representations enriched with fine-grained semantic understanding.Evaluated on six echocardiography datasets spanning fetal, pediatric, and adult populations, DISCOVR outperforms both specialized video anomaly detection methods and state-of-the-art video-SSL baselines in zero-shot and linear probing setups,achieving superior segmentation transfer and strong downstream performance on clinically relevant tasks such as LVEF prediction. Code available at: https://github.com/mdivyanshu97/DISCOVR
title Self-supervised Learning of Echocardiographic Video Representations via Online Cluster Distillation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2506.11777