Enhancing Multimodal Medical Image Classification using Cross-Graph Modal Contrastive Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ding, Jun-En, Hsu, Chien-Chin, Chu, Chi-Hsiang, Wang, Shuqiang, Liu, Feng
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912696676384768
author Ding, Jun-En
Hsu, Chien-Chin
Chu, Chi-Hsiang
Wang, Shuqiang
Liu, Feng
author_facet Ding, Jun-En
Hsu, Chien-Chin
Chu, Chi-Hsiang
Wang, Shuqiang
Liu, Feng
contents The classification of medical images is a pivotal aspect of disease diagnosis, often enhanced by deep learning techniques. However, traditional approaches typically focus on unimodal medical image data, neglecting the integration of diverse non-image patient data. This paper proposes a novel Cross-Graph Modal Contrastive Learning (CGMCL) framework for multimodal structured data from different data domains to improve medical image classification. The model effectively integrates both image and non-image data by constructing cross-modality graphs and leveraging contrastive learning to align multimodal features in a shared latent space. An inter-modality feature scaling module further optimizes the representation learning process by reducing the gap between heterogeneous modalities. The proposed approach is evaluated on two datasets: a Parkinson's disease (PD) dataset and a public melanoma dataset. Results demonstrate that CGMCL outperforms conventional unimodal methods in accuracy, interpretability, and early disease prediction. Additionally, the method shows superior performance in multi-class melanoma classification. The CGMCL framework provides valuable insights into medical image classification while offering improved disease interpretability and predictive capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17494
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Multimodal Medical Image Classification using Cross-Graph Modal Contrastive Learning
Ding, Jun-En
Hsu, Chien-Chin
Chu, Chi-Hsiang
Wang, Shuqiang
Liu, Feng
Image and Video Processing
Computer Vision and Pattern Recognition
The classification of medical images is a pivotal aspect of disease diagnosis, often enhanced by deep learning techniques. However, traditional approaches typically focus on unimodal medical image data, neglecting the integration of diverse non-image patient data. This paper proposes a novel Cross-Graph Modal Contrastive Learning (CGMCL) framework for multimodal structured data from different data domains to improve medical image classification. The model effectively integrates both image and non-image data by constructing cross-modality graphs and leveraging contrastive learning to align multimodal features in a shared latent space. An inter-modality feature scaling module further optimizes the representation learning process by reducing the gap between heterogeneous modalities. The proposed approach is evaluated on two datasets: a Parkinson's disease (PD) dataset and a public melanoma dataset. Results demonstrate that CGMCL outperforms conventional unimodal methods in accuracy, interpretability, and early disease prediction. Additionally, the method shows superior performance in multi-class melanoma classification. The CGMCL framework provides valuable insights into medical image classification while offering improved disease interpretability and predictive capabilities.
title Enhancing Multimodal Medical Image Classification using Cross-Graph Modal Contrastive Learning
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.17494