DCMM-Transformer: Degree-Corrected Mixed-Membership Attention for Medical Imaging
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866909903857123328 |
|---|---|
| author | Cheng, Huimin Yu, Xiaowei Wu, Shushan Fang, Luyang Cao, Chao Zhang, Jing Liu, Tianming Zhu, Dajiang Zhong, Wenxuan Ma, Ping |
| author_facet | Cheng, Huimin Yu, Xiaowei Wu, Shushan Fang, Luyang Cao, Chao Zhang, Jing Liu, Tianming Zhu, Dajiang Zhong, Wenxuan Ma, Ping |
| contents | Medical images exhibit latent anatomical groupings, such as organs, tissues, and pathological regions, that standard Vision Transformers (ViTs) fail to exploit. While recent work like SBM-Transformer attempts to incorporate such structures through stochastic binary masking, they suffer from non-differentiability, training instability, and the inability to model complex community structure. We present DCMM-Transformer, a novel ViT architecture for medical image analysis that incorporates a Degree-Corrected Mixed-Membership (DCMM) model as an additive bias in self-attention. Unlike prior approaches that rely on multiplicative masking and binary sampling, our method introduces community structure and degree heterogeneity in a fully differentiable and interpretable manner. Comprehensive experiments across diverse medical imaging datasets, including brain, chest, breast, and ocular modalities, demonstrate the superior performance and generalizability of the proposed approach. Furthermore, the learned group structure and structured attention modulation substantially enhance interpretability by yielding attention maps that are anatomically meaningful and semantically coherent. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_12047 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DCMM-Transformer: Degree-Corrected Mixed-Membership Attention for Medical Imaging Cheng, Huimin Yu, Xiaowei Wu, Shushan Fang, Luyang Cao, Chao Zhang, Jing Liu, Tianming Zhu, Dajiang Zhong, Wenxuan Ma, Ping Computer Vision and Pattern Recognition Artificial Intelligence Medical images exhibit latent anatomical groupings, such as organs, tissues, and pathological regions, that standard Vision Transformers (ViTs) fail to exploit. While recent work like SBM-Transformer attempts to incorporate such structures through stochastic binary masking, they suffer from non-differentiability, training instability, and the inability to model complex community structure. We present DCMM-Transformer, a novel ViT architecture for medical image analysis that incorporates a Degree-Corrected Mixed-Membership (DCMM) model as an additive bias in self-attention. Unlike prior approaches that rely on multiplicative masking and binary sampling, our method introduces community structure and degree heterogeneity in a fully differentiable and interpretable manner. Comprehensive experiments across diverse medical imaging datasets, including brain, chest, breast, and ocular modalities, demonstrate the superior performance and generalizability of the proposed approach. Furthermore, the learned group structure and structured attention modulation substantially enhance interpretability by yielding attention maps that are anatomically meaningful and semantically coherent. |
| title | DCMM-Transformer: Degree-Corrected Mixed-Membership Attention for Medical Imaging |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2511.12047 |