Demystify Mamba in Vision: A Linear Attention Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Dongchen, Wang, Ziyi, Xia, Zhuofan, Han, Yizeng, Pu, Yifan, Ge, Chunjiang, Song, Jun, Song, Shiji, Zheng, Bo, Huang, Gao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging the Divide: Reconsidering Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2024)
by: Han, Dongchen, et al.
Published: (2024)
Agent Attention: On the Integration of Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2023)
by: Han, Dongchen, et al.
Published: (2023)
GSVA: Generalized Segmentation via Multimodal Large Language Models
by: Xia, Zhuofan, et al.
Published: (2023)
by: Xia, Zhuofan, et al.
Published: (2023)
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
by: Pu, Yifan, et al.
Published: (2024)
by: Pu, Yifan, et al.
Published: (2024)
Vision Transformers are Circulant Attention Learners
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
Latency-aware Unified Dynamic Networks for Efficient Image Recognition
by: Han, Yizeng, et al.
Published: (2023)
by: Han, Yizeng, et al.
Published: (2023)
Linear-Time Global Visual Modeling without Explicit Attention
by: He, Ruize, et al.
Published: (2026)
by: He, Ruize, et al.
Published: (2026)
ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models
by: Ge, Chunjiang, et al.
Published: (2024)
by: Ge, Chunjiang, et al.
Published: (2024)
DyFADet: Dynamic Feature Aggregation for Temporal Action Detection
by: Yang, Le, et al.
Published: (2024)
by: Yang, Le, et al.
Published: (2024)
Cross-Modal Adapter for Vision-Language Retrieval
by: Jiang, Haojun, et al.
Published: (2022)
by: Jiang, Haojun, et al.
Published: (2022)
GRA: Detecting Oriented Objects through Group-wise Rotating and Attention
by: Wang, Jiangshan, et al.
Published: (2024)
by: Wang, Jiangshan, et al.
Published: (2024)
Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
by: Pu, Yifan, et al.
Published: (2025)
by: Pu, Yifan, et al.
Published: (2025)
EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
by: Wang, Yulin, et al.
Published: (2024)
by: Wang, Yulin, et al.
Published: (2024)
Emulating Human-like Adaptive Vision for Efficient and Flexible Machine Visual Perception
by: Wang, Yulin, et al.
Published: (2025)
by: Wang, Yulin, et al.
Published: (2025)
ViT$^3$: Unlocking Test-Time Training in Vision
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
Step by Step Network
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
Exploring contextual modeling with linear complexity for point cloud segmentation
by: Chng, Yong Xien, et al.
Published: (2024)
by: Chng, Yong Xien, et al.
Published: (2024)
Few-Step Distillation for Text-to-Image Generation: A Practical Guide
by: Pu, Yifan, et al.
Published: (2025)
by: Pu, Yifan, et al.
Published: (2025)
SimPro: A Simple Probabilistic Framework Towards Realistic Long-Tailed Semi-Supervised Learning
by: Du, Chaoqun, et al.
Published: (2024)
by: Du, Chaoqun, et al.
Published: (2024)
Everything to the Synthetic: Diffusion-driven Test-time Adaptation via Synthetic-Domain Alignment
by: Guo, Jiayi, et al.
Published: (2024)
by: Guo, Jiayi, et al.
Published: (2024)
Mask Grounding for Referring Image Segmentation
by: Chng, Yong Xien, et al.
Published: (2023)
by: Chng, Yong Xien, et al.
Published: (2023)
Adapting Vision-Language Model with Fine-grained Semantics for Open-Vocabulary Segmentation
by: Chng, Yong Xien, et al.
Published: (2024)
by: Chng, Yong Xien, et al.
Published: (2024)
Training-free Token Reduction for Vision Mamba
by: Ma, Qiankun, et al.
Published: (2025)
by: Ma, Qiankun, et al.
Published: (2025)
EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction
by: Cai, Han, et al.
Published: (2022)
by: Cai, Han, et al.
Published: (2022)
Dynamic Diffusion Transformer
by: Zhao, Wangbo, et al.
Published: (2024)
by: Zhao, Wangbo, et al.
Published: (2024)
Dynamic Tuning Towards Parameter and Inference Efficiency for ViT Adaptation
by: Zhao, Wangbo, et al.
Published: (2024)
by: Zhao, Wangbo, et al.
Published: (2024)
FD-Vision Mamba for Endoscopic Exposure Correction
by: Zheng, Zhuoran, et al.
Published: (2024)
by: Zheng, Zhuoran, et al.
Published: (2024)
Probabilistic Contrastive Learning for Long-Tailed Visual Recognition
by: Du, Chaoqun, et al.
Published: (2024)
by: Du, Chaoqun, et al.
Published: (2024)
Meta-Semi: A Meta-learning Approach for Semi-supervised Learning
by: Wang, Yulin, et al.
Published: (2020)
by: Wang, Yulin, et al.
Published: (2020)
Rethinking the Architecture Design for Efficient Generic Event Boundary Detection
by: Zheng, Ziwei, et al.
Published: (2024)
by: Zheng, Ziwei, et al.
Published: (2024)
DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation
by: Zhao, Wangbo, et al.
Published: (2025)
by: Zhao, Wangbo, et al.
Published: (2025)
The Linear Attention Resurrection in Vision Transformer
by: Zheng, Chuanyang
Published: (2025)
by: Zheng, Chuanyang
Published: (2025)
NeuroMamba: Multi-Perspective Feature Interaction with Visual Mamba for Neuron Segmentation
by: Jiang, Liuyun, et al.
Published: (2026)
by: Jiang, Liuyun, et al.
Published: (2026)
Training an Open-Vocabulary Monocular 3D Object Detection Model without 3D Data
by: Huang, Rui, et al.
Published: (2024)
by: Huang, Rui, et al.
Published: (2024)
Matten: Video Generation with Mamba-Attention
by: Gao, Yu, et al.
Published: (2024)
by: Gao, Yu, et al.
Published: (2024)
UniTTA: Unified Benchmark and Versatile Framework Towards Realistic Test-Time Adaptation
by: Du, Chaoqun, et al.
Published: (2024)
by: Du, Chaoqun, et al.
Published: (2024)
Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
by: Weng, Xingxing, et al.
Published: (2025)
by: Weng, Xingxing, et al.
Published: (2025)
VM-BHINet:Vision Mamba Bimanual Hand Interaction Network for 3D Interacting Hand Mesh Recovery From a Single RGB Image
by: Bi, Han, et al.
Published: (2025)
by: Bi, Han, et al.
Published: (2025)
GlobalMamba: Global Image Serialization for Vision Mamba
by: Wang, Chengkun, et al.
Published: (2024)
by: Wang, Chengkun, et al.
Published: (2024)
CLIP-Mamba: CLIP Pretrained Mamba Models with OOD and Hessian Evaluation
by: Huang, Weiquan, et al.
Published: (2024)
by: Huang, Weiquan, et al.
Published: (2024)
Similar Items
-
Bridging the Divide: Reconsidering Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2024) -
Agent Attention: On the Integration of Softmax and Linear Attention
by: Han, Dongchen, et al.
Published: (2023) -
GSVA: Generalized Segmentation via Multimodal Large Language Models
by: Xia, Zhuofan, et al.
Published: (2023) -
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
by: Pu, Yifan, et al.
Published: (2024) -
Vision Transformers are Circulant Attention Learners
by: Han, Dongchen, et al.
Published: (2025)