ACM-UNet: Adaptive Integration of CNNs and Mamba for Efficient Medical Image Segmentation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Jing, Zhao, Yongkang, Li, Yuhan, Dai, Zhitao, Chen, Cheng, Lai, Qiying
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915313548787712
author Huang, Jing
Zhao, Yongkang
Li, Yuhan
Dai, Zhitao
Chen, Cheng
Lai, Qiying
author_facet Huang, Jing
Zhao, Yongkang
Li, Yuhan
Dai, Zhitao
Chen, Cheng
Lai, Qiying
contents The U-shaped encoder-decoder architecture with skip connections has become a prevailing paradigm in medical image segmentation due to its simplicity and effectiveness. While many recent works aim to improve this framework by designing more powerful encoders and decoders, employing advanced convolutional neural networks (CNNs) for local feature extraction, Transformers or state space models (SSMs) such as Mamba for global context modeling, or hybrid combinations of both, these methods often struggle to fully utilize pretrained vision backbones (e.g., ResNet, ViT, VMamba) due to structural mismatches. To bridge this gap, we introduce ACM-UNet, a general-purpose segmentation framework that retains a simple UNet-like design while effectively incorporating pretrained CNNs and Mamba models through a lightweight adapter mechanism. This adapter resolves architectural incompatibilities and enables the model to harness the complementary strengths of CNNs and SSMs-namely, fine-grained local detail extraction and long-range dependency modeling. Additionally, we propose a hierarchical multi-scale wavelet transform module in the decoder to enhance feature fusion and reconstruction fidelity. Extensive experiments on the Synapse and ACDC benchmarks demonstrate that ACM-UNet achieves state-of-the-art performance while remaining computationally efficient. Notably, it reaches 85.12% Dice Score and 13.89mm HD95 on the Synapse dataset with 17.93G FLOPs, showcasing its effectiveness and scalability. Code is available at: https://github.com/zyklcode/ACM-UNet.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24481
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ACM-UNet: Adaptive Integration of CNNs and Mamba for Efficient Medical Image Segmentation
Huang, Jing
Zhao, Yongkang
Li, Yuhan
Dai, Zhitao
Chen, Cheng
Lai, Qiying
Computer Vision and Pattern Recognition
The U-shaped encoder-decoder architecture with skip connections has become a prevailing paradigm in medical image segmentation due to its simplicity and effectiveness. While many recent works aim to improve this framework by designing more powerful encoders and decoders, employing advanced convolutional neural networks (CNNs) for local feature extraction, Transformers or state space models (SSMs) such as Mamba for global context modeling, or hybrid combinations of both, these methods often struggle to fully utilize pretrained vision backbones (e.g., ResNet, ViT, VMamba) due to structural mismatches. To bridge this gap, we introduce ACM-UNet, a general-purpose segmentation framework that retains a simple UNet-like design while effectively incorporating pretrained CNNs and Mamba models through a lightweight adapter mechanism. This adapter resolves architectural incompatibilities and enables the model to harness the complementary strengths of CNNs and SSMs-namely, fine-grained local detail extraction and long-range dependency modeling. Additionally, we propose a hierarchical multi-scale wavelet transform module in the decoder to enhance feature fusion and reconstruction fidelity. Extensive experiments on the Synapse and ACDC benchmarks demonstrate that ACM-UNet achieves state-of-the-art performance while remaining computationally efficient. Notably, it reaches 85.12% Dice Score and 13.89mm HD95 on the Synapse dataset with 17.93G FLOPs, showcasing its effectiveness and scalability. Code is available at: https://github.com/zyklcode/ACM-UNet.
title ACM-UNet: Adaptive Integration of CNNs and Mamba for Efficient Medical Image Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.24481