UAM: A Unified Attention-Mamba Backbone of Multimodal Framework for Tumor Cell Classification

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Taixi, Chen, Jingyun, Guo, Nancy
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914372610162688
author Chen, Taixi
Chen, Jingyun
Guo, Nancy
author_facet Chen, Taixi
Chen, Jingyun
Guo, Nancy
contents Inspired by the recent success of the Mamba architecture in vision and language domains, we introduce a Unified Attention-Mamba (UAM) backbone. Unlike previous hybrid approaches that integrate Attention and Mamba modules in fixed proportions, our unified design flexibly combines their capabilities within a single cohesive architecture, eliminating the need for manual ratio tuning and improving encode capability. We develop two UAM variants to comprehensively evaluate the benefits of this unified structure. Building on this backbone, we further propose a multimodal UAM framework that jointly performs cell-level classification and image segmentation. Experimental results demonstrate that UAM achieves state-of-the-art performance across both tasks on public benchmarks, surpassing leading image-based foundation models. It improves cell classification accuracy from 74\% to 78\% ($n$=349,882 cells), and tumor segmentation precision from 75\% to 80\% ($n$=406 patches).
format Preprint
id arxiv_https___arxiv_org_abs_2511_17355
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UAM: A Unified Attention-Mamba Backbone of Multimodal Framework for Tumor Cell Classification
Chen, Taixi
Chen, Jingyun
Guo, Nancy
Computer Vision and Pattern Recognition
Inspired by the recent success of the Mamba architecture in vision and language domains, we introduce a Unified Attention-Mamba (UAM) backbone. Unlike previous hybrid approaches that integrate Attention and Mamba modules in fixed proportions, our unified design flexibly combines their capabilities within a single cohesive architecture, eliminating the need for manual ratio tuning and improving encode capability. We develop two UAM variants to comprehensively evaluate the benefits of this unified structure. Building on this backbone, we further propose a multimodal UAM framework that jointly performs cell-level classification and image segmentation. Experimental results demonstrate that UAM achieves state-of-the-art performance across both tasks on public benchmarks, surpassing leading image-based foundation models. It improves cell classification accuracy from 74\% to 78\% ($n$=349,882 cells), and tumor segmentation precision from 75\% to 80\% ($n$=406 patches).
title UAM: A Unified Attention-Mamba Backbone of Multimodal Framework for Tumor Cell Classification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.17355