DCMM-Transformer: Degree-Corrected Mixed-Membership Attention for Medical Imaging

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cheng, Huimin, Yu, Xiaowei, Wu, Shushan, Fang, Luyang, Cao, Chao, Zhang, Jing, Liu, Tianming, Zhu, Dajiang, Zhong, Wenxuan, Ma, Ping
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909903857123328
author Cheng, Huimin
Yu, Xiaowei
Wu, Shushan
Fang, Luyang
Cao, Chao
Zhang, Jing
Liu, Tianming
Zhu, Dajiang
Zhong, Wenxuan
Ma, Ping
author_facet Cheng, Huimin
Yu, Xiaowei
Wu, Shushan
Fang, Luyang
Cao, Chao
Zhang, Jing
Liu, Tianming
Zhu, Dajiang
Zhong, Wenxuan
Ma, Ping
contents Medical images exhibit latent anatomical groupings, such as organs, tissues, and pathological regions, that standard Vision Transformers (ViTs) fail to exploit. While recent work like SBM-Transformer attempts to incorporate such structures through stochastic binary masking, they suffer from non-differentiability, training instability, and the inability to model complex community structure. We present DCMM-Transformer, a novel ViT architecture for medical image analysis that incorporates a Degree-Corrected Mixed-Membership (DCMM) model as an additive bias in self-attention. Unlike prior approaches that rely on multiplicative masking and binary sampling, our method introduces community structure and degree heterogeneity in a fully differentiable and interpretable manner. Comprehensive experiments across diverse medical imaging datasets, including brain, chest, breast, and ocular modalities, demonstrate the superior performance and generalizability of the proposed approach. Furthermore, the learned group structure and structured attention modulation substantially enhance interpretability by yielding attention maps that are anatomically meaningful and semantically coherent.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12047
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DCMM-Transformer: Degree-Corrected Mixed-Membership Attention for Medical Imaging
Cheng, Huimin
Yu, Xiaowei
Wu, Shushan
Fang, Luyang
Cao, Chao
Zhang, Jing
Liu, Tianming
Zhu, Dajiang
Zhong, Wenxuan
Ma, Ping
Computer Vision and Pattern Recognition
Artificial Intelligence
Medical images exhibit latent anatomical groupings, such as organs, tissues, and pathological regions, that standard Vision Transformers (ViTs) fail to exploit. While recent work like SBM-Transformer attempts to incorporate such structures through stochastic binary masking, they suffer from non-differentiability, training instability, and the inability to model complex community structure. We present DCMM-Transformer, a novel ViT architecture for medical image analysis that incorporates a Degree-Corrected Mixed-Membership (DCMM) model as an additive bias in self-attention. Unlike prior approaches that rely on multiplicative masking and binary sampling, our method introduces community structure and degree heterogeneity in a fully differentiable and interpretable manner. Comprehensive experiments across diverse medical imaging datasets, including brain, chest, breast, and ocular modalities, demonstrate the superior performance and generalizability of the proposed approach. Furthermore, the learned group structure and structured attention modulation substantially enhance interpretability by yielding attention maps that are anatomically meaningful and semantically coherent.
title DCMM-Transformer: Degree-Corrected Mixed-Membership Attention for Medical Imaging
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.12047