Self-supervised learning of imaging and clinical signatures using a multimodal joint-embedding predictive architecture

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Thomas Z., Krishnan, Aravind R., Zuo, Lianrui, Still, John M., Sandler, Kim L., Maldonado, Fabien, Lasko, Thomas A., Landman, Bennett A.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914046948671488
author Li, Thomas Z.
Krishnan, Aravind R.
Zuo, Lianrui
Still, John M.
Sandler, Kim L.
Maldonado, Fabien
Lasko, Thomas A.
Landman, Bennett A.
author_facet Li, Thomas Z.
Krishnan, Aravind R.
Zuo, Lianrui
Still, John M.
Sandler, Kim L.
Maldonado, Fabien
Lasko, Thomas A.
Landman, Bennett A.
contents The development of multimodal models for pulmonary nodule diagnosis is limited by the scarcity of labeled data and the tendency for these models to overfit on the training distribution. In this work, we leverage self-supervised learning from longitudinal and multimodal archives to address these challenges. We curate an unlabeled set of patients with CT scans and linked electronic health records from our home institution to power joint embedding predictive architecture (JEPA) pretraining. After supervised finetuning, we show that our approach outperforms an unregularized multimodal model and imaging-only model in an internal cohort (ours: 0.91, multimodal: 0.88, imaging-only: 0.73 AUC), but underperforms in an external cohort (ours: 0.72, imaging-only: 0.75 AUC). We develop a synthetic environment that characterizes the context in which JEPA may underperform. This work innovates an approach that leverages unlabeled multimodal medical archives to improve predictive models and demonstrates its advantages and limitations in pulmonary nodule diagnosis.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15470
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Self-supervised learning of imaging and clinical signatures using a multimodal joint-embedding predictive architecture
Li, Thomas Z.
Krishnan, Aravind R.
Zuo, Lianrui
Still, John M.
Sandler, Kim L.
Maldonado, Fabien
Lasko, Thomas A.
Landman, Bennett A.
Computer Vision and Pattern Recognition
Artificial Intelligence
The development of multimodal models for pulmonary nodule diagnosis is limited by the scarcity of labeled data and the tendency for these models to overfit on the training distribution. In this work, we leverage self-supervised learning from longitudinal and multimodal archives to address these challenges. We curate an unlabeled set of patients with CT scans and linked electronic health records from our home institution to power joint embedding predictive architecture (JEPA) pretraining. After supervised finetuning, we show that our approach outperforms an unregularized multimodal model and imaging-only model in an internal cohort (ours: 0.91, multimodal: 0.88, imaging-only: 0.73 AUC), but underperforms in an external cohort (ours: 0.72, imaging-only: 0.75 AUC). We develop a synthetic environment that characterizes the context in which JEPA may underperform. This work innovates an approach that leverages unlabeled multimodal medical archives to improve predictive models and demonstrates its advantages and limitations in pulmonary nodule diagnosis.
title Self-supervised learning of imaging and clinical signatures using a multimodal joint-embedding predictive architecture
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.15470