Data distribution impacts the performance and generalisability of contrastive learning-based foundation models of electrocardiograms

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Khattak, Gul Rukh, Patlatzoglou, Konstantinos, Barker, Joseph, Pastika, Libor, Zeidaabadi, Boroumand, El-Medany, Ahmed, Aggour, Hesham, Liang, Yixiu, Ribeiro, Antonio H., Annis, Jeffrey, Ribeiro, Antonio Luiz Pinho, Ge, Junbo, Kramer, Daniel B., Waks, Jonathan W., Brittain, Evan, Peters, Nicholas, Ng, Fu Siong, Sau, Arunashis
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911151516811264
author Khattak, Gul Rukh
Patlatzoglou, Konstantinos
Barker, Joseph
Pastika, Libor
Zeidaabadi, Boroumand
El-Medany, Ahmed
Aggour, Hesham
Liang, Yixiu
Ribeiro, Antonio H.
Annis, Jeffrey
Ribeiro, Antonio Luiz Pinho
Ge, Junbo
Kramer, Daniel B.
Waks, Jonathan W.
Brittain, Evan
Peters, Nicholas
Ng, Fu Siong
Sau, Arunashis
author_facet Khattak, Gul Rukh
Patlatzoglou, Konstantinos
Barker, Joseph
Pastika, Libor
Zeidaabadi, Boroumand
El-Medany, Ahmed
Aggour, Hesham
Liang, Yixiu
Ribeiro, Antonio H.
Annis, Jeffrey
Ribeiro, Antonio Luiz Pinho
Ge, Junbo
Kramer, Daniel B.
Waks, Jonathan W.
Brittain, Evan
Peters, Nicholas
Ng, Fu Siong
Sau, Arunashis
contents Contrastive learning is a widely adopted self-supervised pretraining strategy, yet its dependence on cohort composition remains underexplored. We present Contrasting by Patient Augmented Electrocardiograms (CAPE) foundation model and pretrain on four cohorts (n = 5,203,352), from diverse populations across three continents (North America, South America, Asia). We systematically assess how cohort demographics, health status, and population diversity influence the downstream performance for prediction tasks also including two additional cohorts from another continent (Europe). We find that downstream performance depends on the distributional properties of the pretraining cohort, including demographics and health status. Moreover, while pretraining with a multi-centre, demographically diverse cohort improves in-distribution accuracy, it reduces out-of-distribution (OOD) generalisation of our contrastive approach by encoding cohort-specific artifacts. To address this, we propose the In-Distribution Batch (IDB) strategy, which preserves intra-cohort consistency during pretraining and enhances OOD robustness. This work provides important insights for developing clinically fair and generalisable foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_10369
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Data distribution impacts the performance and generalisability of contrastive learning-based foundation models of electrocardiograms
Khattak, Gul Rukh
Patlatzoglou, Konstantinos
Barker, Joseph
Pastika, Libor
Zeidaabadi, Boroumand
El-Medany, Ahmed
Aggour, Hesham
Liang, Yixiu
Ribeiro, Antonio H.
Annis, Jeffrey
Ribeiro, Antonio Luiz Pinho
Ge, Junbo
Kramer, Daniel B.
Waks, Jonathan W.
Brittain, Evan
Peters, Nicholas
Ng, Fu Siong
Sau, Arunashis
Machine Learning
Artificial Intelligence
Signal Processing
Tissues and Organs
Contrastive learning is a widely adopted self-supervised pretraining strategy, yet its dependence on cohort composition remains underexplored. We present Contrasting by Patient Augmented Electrocardiograms (CAPE) foundation model and pretrain on four cohorts (n = 5,203,352), from diverse populations across three continents (North America, South America, Asia). We systematically assess how cohort demographics, health status, and population diversity influence the downstream performance for prediction tasks also including two additional cohorts from another continent (Europe). We find that downstream performance depends on the distributional properties of the pretraining cohort, including demographics and health status. Moreover, while pretraining with a multi-centre, demographically diverse cohort improves in-distribution accuracy, it reduces out-of-distribution (OOD) generalisation of our contrastive approach by encoding cohort-specific artifacts. To address this, we propose the In-Distribution Batch (IDB) strategy, which preserves intra-cohort consistency during pretraining and enhances OOD robustness. This work provides important insights for developing clinically fair and generalisable foundation models.
title Data distribution impacts the performance and generalisability of contrastive learning-based foundation models of electrocardiograms
topic Machine Learning
Artificial Intelligence
Signal Processing
Tissues and Organs
url https://arxiv.org/abs/2509.10369