MultiViT2: A Data-augmented Multimodal Neuroimaging Prediction Framework via Latent Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuda, Bi, Sihan, Jia, Yutong, Gao, Anees, Abrol, Zening, Fu, Vince, Calhoun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911007549423616
author Yuda, Bi
Sihan, Jia
Yutong, Gao
Anees, Abrol
Zening, Fu
Vince, Calhoun
author_facet Yuda, Bi
Sihan, Jia
Yutong, Gao
Anees, Abrol
Zening, Fu
Vince, Calhoun
contents Multimodal medical imaging integrates diverse data types, such as structural and functional neuroimaging, to provide complementary insights that enhance deep learning predictions and improve outcomes. This study focuses on a neuroimaging prediction framework based on both structural and functional neuroimaging data. We propose a next-generation prediction model, \textbf{MultiViT2}, which combines a pretrained representative learning base model with a vision transformer backbone for prediction output. Additionally, we developed a data augmentation module based on the latent diffusion model that enriches input data by generating augmented neuroimaging samples, thereby enhancing predictive performance through reduced overfitting and improved generalizability. We show that MultiViT2 significantly outperforms the first-generation model in schizophrenia classification accuracy and demonstrates strong scalability and portability.
format Preprint
id arxiv_https___arxiv_org_abs_2506_13667
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MultiViT2: A Data-augmented Multimodal Neuroimaging Prediction Framework via Latent Diffusion Model
Yuda, Bi
Sihan, Jia
Yutong, Gao
Anees, Abrol
Zening, Fu
Vince, Calhoun
Image and Video Processing
Computer Vision and Pattern Recognition
Multimodal medical imaging integrates diverse data types, such as structural and functional neuroimaging, to provide complementary insights that enhance deep learning predictions and improve outcomes. This study focuses on a neuroimaging prediction framework based on both structural and functional neuroimaging data. We propose a next-generation prediction model, \textbf{MultiViT2}, which combines a pretrained representative learning base model with a vision transformer backbone for prediction output. Additionally, we developed a data augmentation module based on the latent diffusion model that enriches input data by generating augmented neuroimaging samples, thereby enhancing predictive performance through reduced overfitting and improved generalizability. We show that MultiViT2 significantly outperforms the first-generation model in schizophrenia classification accuracy and demonstrates strong scalability and portability.
title MultiViT2: A Data-augmented Multimodal Neuroimaging Prediction Framework via Latent Diffusion Model
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.13667