Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.13858 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916662695952384 |
|---|---|
| author | Ke, Hongyu Morris, Jack Oguchi, Kentaro Cao, Xiaofei Liu, Yongkang Wang, Haoxin Ding, Yi |
| author_facet | Ke, Hongyu Morris, Jack Oguchi, Kentaro Cao, Xiaofei Liu, Yongkang Wang, Haoxin Ding, Yi |
| contents | 3D visual perception tasks, such as 3D detection from multi-camera images, are essential components of autonomous driving and assistance systems. However, designing computationally efficient methods remains a significant challenge. In this paper, we propose a Mamba-based framework called MamBEV, which learns unified Bird's Eye View (BEV) representations using linear spatio-temporal SSM-based attention. This approach supports multiple 3D perception tasks with significantly improved computational and memory efficiency. Furthermore, we introduce SSM based cross-attention, analogous to standard cross attention, where BEV query representations can interact with relevant image features. Extensive experiments demonstrate MamBEV's promising performance across diverse visual perception metrics, highlighting its advantages in input scaling efficiency compared to existing benchmark models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_13858 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MamBEV: Enabling State Space Models to Learn Birds-Eye-View Representations Ke, Hongyu Morris, Jack Oguchi, Kentaro Cao, Xiaofei Liu, Yongkang Wang, Haoxin Ding, Yi Computer Vision and Pattern Recognition Machine Learning 3D visual perception tasks, such as 3D detection from multi-camera images, are essential components of autonomous driving and assistance systems. However, designing computationally efficient methods remains a significant challenge. In this paper, we propose a Mamba-based framework called MamBEV, which learns unified Bird's Eye View (BEV) representations using linear spatio-temporal SSM-based attention. This approach supports multiple 3D perception tasks with significantly improved computational and memory efficiency. Furthermore, we introduce SSM based cross-attention, analogous to standard cross attention, where BEV query representations can interact with relevant image features. Extensive experiments demonstrate MamBEV's promising performance across diverse visual perception metrics, highlighting its advantages in input scaling efficiency compared to existing benchmark models. |
| title | MamBEV: Enabling State Space Models to Learn Birds-Eye-View Representations |
| topic | Computer Vision and Pattern Recognition Machine Learning |
| url | https://arxiv.org/abs/2503.13858 |