Saved in:
Bibliographic Details
Main Authors: Ke, Hongyu, Morris, Jack, Oguchi, Kentaro, Cao, Xiaofei, Liu, Yongkang, Wang, Haoxin, Ding, Yi
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2503.13858
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916662695952384
author Ke, Hongyu
Morris, Jack
Oguchi, Kentaro
Cao, Xiaofei
Liu, Yongkang
Wang, Haoxin
Ding, Yi
author_facet Ke, Hongyu
Morris, Jack
Oguchi, Kentaro
Cao, Xiaofei
Liu, Yongkang
Wang, Haoxin
Ding, Yi
contents 3D visual perception tasks, such as 3D detection from multi-camera images, are essential components of autonomous driving and assistance systems. However, designing computationally efficient methods remains a significant challenge. In this paper, we propose a Mamba-based framework called MamBEV, which learns unified Bird's Eye View (BEV) representations using linear spatio-temporal SSM-based attention. This approach supports multiple 3D perception tasks with significantly improved computational and memory efficiency. Furthermore, we introduce SSM based cross-attention, analogous to standard cross attention, where BEV query representations can interact with relevant image features. Extensive experiments demonstrate MamBEV's promising performance across diverse visual perception metrics, highlighting its advantages in input scaling efficiency compared to existing benchmark models.
format Preprint
id arxiv_https___arxiv_org_abs_2503_13858
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MamBEV: Enabling State Space Models to Learn Birds-Eye-View Representations
Ke, Hongyu
Morris, Jack
Oguchi, Kentaro
Cao, Xiaofei
Liu, Yongkang
Wang, Haoxin
Ding, Yi
Computer Vision and Pattern Recognition
Machine Learning
3D visual perception tasks, such as 3D detection from multi-camera images, are essential components of autonomous driving and assistance systems. However, designing computationally efficient methods remains a significant challenge. In this paper, we propose a Mamba-based framework called MamBEV, which learns unified Bird's Eye View (BEV) representations using linear spatio-temporal SSM-based attention. This approach supports multiple 3D perception tasks with significantly improved computational and memory efficiency. Furthermore, we introduce SSM based cross-attention, analogous to standard cross attention, where BEV query representations can interact with relevant image features. Extensive experiments demonstrate MamBEV's promising performance across diverse visual perception metrics, highlighting its advantages in input scaling efficiency compared to existing benchmark models.
title MamBEV: Enabling State Space Models to Learn Birds-Eye-View Representations
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2503.13858