Multimodal Visual Surrogate Compression for Alzheimer's Disease Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Dexuan, Peng, Ciyuan, Kuantama, Endrowednes, Guo, Jingcai, Wu, Jia, Yang, Jian, Beheshti, Amin, Yang, Ming-Hsuan, Qi, Yuankai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917232238395392
author Ding, Dexuan
Peng, Ciyuan
Kuantama, Endrowednes
Guo, Jingcai
Wu, Jia
Yang, Jian
Beheshti, Amin
Yang, Ming-Hsuan
Qi, Yuankai
author_facet Ding, Dexuan
Peng, Ciyuan
Kuantama, Endrowednes
Guo, Jingcai
Wu, Jia
Yang, Jian
Beheshti, Amin
Yang, Ming-Hsuan
Qi, Yuankai
contents High-dimensional structural MRI (sMRI) images are widely used for Alzheimer's Disease (AD) diagnosis. Most existing methods for sMRI representation learning rely on 3D architectures (e.g., 3D CNNs), slice-wise feature extraction with late aggregation, or apply training-free feature extractions using 2D foundation models (e.g., DINO). However, these three paradigms suffer from high computational cost, loss of cross-slice relations, and limited ability to extract discriminative features, respectively. To address these challenges, we propose Multimodal Visual Surrogate Compression (MVSC). It learns to compress and adapt large 3D sMRI volumes into compact 2D features, termed as visual surrogates, which are better aligned with frozen 2D foundation models to extract powerful representations for final AD classification. MVSC has two key components: a Volume Context Encoder that captures global cross-slice context under textual guidance, and an Adaptive Slice Fusion module that aggregates slice-level information in a text-enhanced, patch-wise manner. Extensive experiments on three large-scale Alzheimer's disease benchmarks demonstrate our MVSC performs favourably on both binary and multi-class classification tasks compared against state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2601_21673
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Multimodal Visual Surrogate Compression for Alzheimer's Disease Classification
Ding, Dexuan
Peng, Ciyuan
Kuantama, Endrowednes
Guo, Jingcai
Wu, Jia
Yang, Jian
Beheshti, Amin
Yang, Ming-Hsuan
Qi, Yuankai
Computer Vision and Pattern Recognition
High-dimensional structural MRI (sMRI) images are widely used for Alzheimer's Disease (AD) diagnosis. Most existing methods for sMRI representation learning rely on 3D architectures (e.g., 3D CNNs), slice-wise feature extraction with late aggregation, or apply training-free feature extractions using 2D foundation models (e.g., DINO). However, these three paradigms suffer from high computational cost, loss of cross-slice relations, and limited ability to extract discriminative features, respectively. To address these challenges, we propose Multimodal Visual Surrogate Compression (MVSC). It learns to compress and adapt large 3D sMRI volumes into compact 2D features, termed as visual surrogates, which are better aligned with frozen 2D foundation models to extract powerful representations for final AD classification. MVSC has two key components: a Volume Context Encoder that captures global cross-slice context under textual guidance, and an Adaptive Slice Fusion module that aggregates slice-level information in a text-enhanced, patch-wise manner. Extensive experiments on three large-scale Alzheimer's disease benchmarks demonstrate our MVSC performs favourably on both binary and multi-class classification tasks compared against state-of-the-art methods.
title Multimodal Visual Surrogate Compression for Alzheimer's Disease Classification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.21673