Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Lianghui, Liao, Bencheng, Zhang, Qian, Wang, Xinlong, Liu, Wenyu, Wang, Xinggang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
von: Zou, Jialv, et al.
Veröffentlicht: (2025)
von: Zou, Jialv, et al.
Veröffentlicht: (2025)
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
von: Li, Yingyue, et al.
Veröffentlicht: (2025)
von: Li, Yingyue, et al.
Veröffentlicht: (2025)
ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention
von: Liao, Bencheng, et al.
Veröffentlicht: (2024)
von: Liao, Bencheng, et al.
Veröffentlicht: (2024)
MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
von: Zou, Jialv, et al.
Veröffentlicht: (2024)
von: Zou, Jialv, et al.
Veröffentlicht: (2024)
Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
von: Liao, Bencheng, et al.
Veröffentlicht: (2025)
von: Liao, Bencheng, et al.
Veröffentlicht: (2025)
InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
von: Tao, Hongyuan, et al.
Veröffentlicht: (2025)
von: Tao, Hongyuan, et al.
Veröffentlicht: (2025)
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
von: Zeng, Lunbin, et al.
Veröffentlicht: (2025)
von: Zeng, Lunbin, et al.
Veröffentlicht: (2025)
DiG: Scalable and Efficient Diffusion Models with Gated Linear Attention
von: Zhu, Lianghui, et al.
Veröffentlicht: (2024)
von: Zhu, Lianghui, et al.
Veröffentlicht: (2024)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
von: Jiang, Bo, et al.
Veröffentlicht: (2024)
von: Jiang, Bo, et al.
Veröffentlicht: (2024)
Surface Vision Mamba: Leveraging Bidirectional State Space Model for Efficient Spherical Manifold Representation
von: He, Rongzhao, et al.
Veröffentlicht: (2025)
von: He, Rongzhao, et al.
Veröffentlicht: (2025)
VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
von: Jiang, Bo, et al.
Veröffentlicht: (2024)
von: Jiang, Bo, et al.
Veröffentlicht: (2024)
DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
von: Zou, Jialv, et al.
Veröffentlicht: (2025)
von: Zou, Jialv, et al.
Veröffentlicht: (2025)
WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
Lane Graph as Path: Continuity-preserving Path-wise Modeling for Online Lane Graph Construction
von: Liao, Bencheng, et al.
Veröffentlicht: (2023)
von: Liao, Bencheng, et al.
Veröffentlicht: (2023)
MapTRv2: An End-to-End Framework for Online Vectorized HD Map Construction
von: Liao, Bencheng, et al.
Veröffentlicht: (2023)
von: Liao, Bencheng, et al.
Veröffentlicht: (2023)
EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuxuan, et al.
Veröffentlicht: (2024)
WeakSAM: Segment Anything Meets Weakly-supervised Instance-level Recognition
von: Zhu, Lianghui, et al.
Veröffentlicht: (2024)
von: Zhu, Lianghui, et al.
Veröffentlicht: (2024)
PersonViT: Large-scale Self-supervised Vision Transformer for Person Re-Identification
von: Hu, Bin, et al.
Veröffentlicht: (2024)
von: Hu, Bin, et al.
Veröffentlicht: (2024)
TransLight: Image-Guided Customized Lighting Control with Generative Decoupling
von: Li, Zongming, et al.
Veröffentlicht: (2025)
von: Li, Zongming, et al.
Veröffentlicht: (2025)
DeltaMIL: Gated Memory Integration for Efficient and Discriminative Whole Slide Image Analysis
von: Zhu, Yueting, et al.
Veröffentlicht: (2025)
von: Zhu, Yueting, et al.
Veröffentlicht: (2025)
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
von: Cheng, Tianheng, et al.
Veröffentlicht: (2026)
von: Cheng, Tianheng, et al.
Veröffentlicht: (2026)
Polar Parametrization for Vision-based Surround-View 3D Detection
von: Chen, Shaoyu, et al.
Veröffentlicht: (2022)
von: Chen, Shaoyu, et al.
Veröffentlicht: (2022)
GaraMoSt: Parallel Multi-Granularity Motion and Structural Modeling for Efficient Multi-Frame Interpolation in DSA Images
von: Xu, Ziyang, et al.
Veröffentlicht: (2024)
von: Xu, Ziyang, et al.
Veröffentlicht: (2024)
EVA-02: A Visual Representation for Neon Genesis
von: Fang, Yuxin, et al.
Veröffentlicht: (2023)
von: Fang, Yuxin, et al.
Veröffentlicht: (2023)
MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning
von: Zhang, Wenrui, et al.
Veröffentlicht: (2025)
von: Zhang, Wenrui, et al.
Veröffentlicht: (2025)
AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
von: Jiang, Bo, et al.
Veröffentlicht: (2025)
DefMamba: Deformable Visual State Space Model
von: Liu, Leiye, et al.
Veröffentlicht: (2025)
von: Liu, Leiye, et al.
Veröffentlicht: (2025)
QuadMamba: Learning Quadtree-based Selective Scan for Visual State Space Model
von: Xie, Fei, et al.
Veröffentlicht: (2024)
von: Xie, Fei, et al.
Veröffentlicht: (2024)
Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained Environments
von: Shihab, Ibne Farabi, et al.
Veröffentlicht: (2025)
von: Shihab, Ibne Farabi, et al.
Veröffentlicht: (2025)
RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework
von: Gao, Hao, et al.
Veröffentlicht: (2026)
von: Gao, Hao, et al.
Veröffentlicht: (2026)
Causality-inspired Discriminative Feature Learning in Triple Domains for Gait Recognition
von: Xiong, Haijun, et al.
Veröffentlicht: (2024)
von: Xiong, Haijun, et al.
Veröffentlicht: (2024)
Understanding Self-Supervised Pretraining with Part-Aware Representation Learning
von: Zhu, Jie, et al.
Veröffentlicht: (2023)
von: Zhu, Jie, et al.
Veröffentlicht: (2023)
VideoMamba: State Space Model for Efficient Video Understanding
von: Li, Kunchang, et al.
Veröffentlicht: (2024)
von: Li, Kunchang, et al.
Veröffentlicht: (2024)
Matte Anything: Interactive Natural Image Matting with Segment Anything Models
von: Yao, Jingfeng, et al.
Veröffentlicht: (2023)
von: Yao, Jingfeng, et al.
Veröffentlicht: (2023)
HTD-Mamba: Efficient Hyperspectral Target Detection with Pyramid State Space Model
von: Shen, Dunbin, et al.
Veröffentlicht: (2024)
von: Shen, Dunbin, et al.
Veröffentlicht: (2024)
GroupMamba: Efficient Group-Based Visual State Space Model
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2024)
GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
GraphMamba: An Efficient Graph Structure Learning Vision Mamba for Hyperspectral Image Classification
von: Yang, Aitao, et al.
Veröffentlicht: (2024)
von: Yang, Aitao, et al.
Veröffentlicht: (2024)
VCMamba: Bridging Convolutions with Multi-Directional Mamba for Efficient Visual Representation
von: Munir, Mustafa, et al.
Veröffentlicht: (2025)
von: Munir, Mustafa, et al.
Veröffentlicht: (2025)
ViTGaze: Gaze Following with Interaction Features in Vision Transformers
von: Song, Yuehao, et al.
Veröffentlicht: (2024)
von: Song, Yuehao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
von: Zou, Jialv, et al.
Veröffentlicht: (2025) -
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
von: Li, Yingyue, et al.
Veröffentlicht: (2025) -
ViG: Linear-complexity Visual Sequence Learning with Gated Linear Attention
von: Liao, Bencheng, et al.
Veröffentlicht: (2024) -
MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
von: Zou, Jialv, et al.
Veröffentlicht: (2024) -
Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
von: Liao, Bencheng, et al.
Veröffentlicht: (2025)