Saved in:
Bibliographic Details
Main Authors: Niu, Ziwei, Sun, Hao, Bian, Shujun, Yang, Xihong, Lin, Lanfen, Liu, Yuxin, Jin, Yueming
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.21154
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912924109373440
author Niu, Ziwei
Sun, Hao
Bian, Shujun
Yang, Xihong
Lin, Lanfen
Liu, Yuxin
Jin, Yueming
author_facet Niu, Ziwei
Sun, Hao
Bian, Shujun
Yang, Xihong
Lin, Lanfen
Liu, Yuxin
Jin, Yueming
contents Accurate interpretation of electrocardiogram (ECG) signals is crucial for diagnosing cardiovascular diseases. Recent multimodal approaches that integrate ECGs with accompanying clinical reports show strong potential, but they still face two main concerns from a modality perspective: (1) intra-modality: existing models process ECGs in a lead-agnostic manner, overlooking spatial-temporal dependencies across leads, which restricts their effectiveness in modeling fine-grained diagnostic patterns; (2) inter-modality: existing methods directly align ECG signals with clinical reports, introducing modality-specific biases due to the free-text nature of the reports. In light of these two issues, we propose CG-DMER, a contrastive-generative framework for disentangled multimodal ECG representation learning, powered by two key designs: (1) Spatial-temporal masked modeling is designed to better capture fine-grained temporal dynamics and inter-lead spatial dependencies by applying masking across both spatial and temporal dimensions and reconstructing the missing information. (2) A representation disentanglement and alignment strategy is designed to mitigate unnecessary noise and modality-specific biases by introducing modality-specific and modality-shared encoders, ensuring a clearer separation between modality-invariant and modality-specific representations. Experiments on three public datasets demonstrate that CG-DMER achieves state-of-the-art performance across diverse downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2602_21154
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CG-DMER: Hybrid Contrastive-Generative Framework for Disentangled Multimodal ECG Representation Learning
Niu, Ziwei
Sun, Hao
Bian, Shujun
Yang, Xihong
Lin, Lanfen
Liu, Yuxin
Jin, Yueming
Artificial Intelligence
Accurate interpretation of electrocardiogram (ECG) signals is crucial for diagnosing cardiovascular diseases. Recent multimodal approaches that integrate ECGs with accompanying clinical reports show strong potential, but they still face two main concerns from a modality perspective: (1) intra-modality: existing models process ECGs in a lead-agnostic manner, overlooking spatial-temporal dependencies across leads, which restricts their effectiveness in modeling fine-grained diagnostic patterns; (2) inter-modality: existing methods directly align ECG signals with clinical reports, introducing modality-specific biases due to the free-text nature of the reports. In light of these two issues, we propose CG-DMER, a contrastive-generative framework for disentangled multimodal ECG representation learning, powered by two key designs: (1) Spatial-temporal masked modeling is designed to better capture fine-grained temporal dynamics and inter-lead spatial dependencies by applying masking across both spatial and temporal dimensions and reconstructing the missing information. (2) A representation disentanglement and alignment strategy is designed to mitigate unnecessary noise and modality-specific biases by introducing modality-specific and modality-shared encoders, ensuring a clearer separation between modality-invariant and modality-specific representations. Experiments on three public datasets demonstrate that CG-DMER achieves state-of-the-art performance across diverse downstream tasks.
title CG-DMER: Hybrid Contrastive-Generative Framework for Disentangled Multimodal ECG Representation Learning
topic Artificial Intelligence
url https://arxiv.org/abs/2602.21154