PointBeV: A Sparse Approach to BeV Predictions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chambon, Loick, Zablocki, Eloi, Chen, Mickael, Bartoccioni, Florent, Perez, Patrick, Cord, Matthieu
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913358923431936
author Chambon, Loick
Zablocki, Eloi
Chen, Mickael
Bartoccioni, Florent
Perez, Patrick
Cord, Matthieu
author_facet Chambon, Loick
Zablocki, Eloi
Chen, Mickael
Bartoccioni, Florent
Perez, Patrick
Cord, Matthieu
contents Bird's-eye View (BeV) representations have emerged as the de-facto shared space in driving applications, offering a unified space for sensor data fusion and supporting various downstream tasks. However, conventional models use grids with fixed resolution and range and face computational inefficiencies due to the uniform allocation of resources across all cells. To address this, we propose PointBeV, a novel sparse BeV segmentation model operating on sparse BeV cells instead of dense grids. This approach offers precise control over memory usage, enabling the use of long temporal contexts and accommodating memory-constrained platforms. PointBeV employs an efficient two-pass strategy for training, enabling focused computation on regions of interest. At inference time, it can be used with various memory/performance trade-offs and flexibly adjusts to new specific use cases. PointBeV achieves state-of-the-art results on the nuScenes dataset for vehicle, pedestrian, and lane segmentation, showcasing superior performance in static and temporal settings despite being trained solely with sparse signals. We will release our code along with two new efficient modules used in the architecture: Sparse Feature Pulling, designed for the effective extraction of features from images to BeV, and Submanifold Attention, which enables efficient temporal modeling. Our code is available at https://github.com/valeoai/PointBeV.
format Preprint
id arxiv_https___arxiv_org_abs_2312_00703
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle PointBeV: A Sparse Approach to BeV Predictions
Chambon, Loick
Zablocki, Eloi
Chen, Mickael
Bartoccioni, Florent
Perez, Patrick
Cord, Matthieu
Computer Vision and Pattern Recognition
Bird's-eye View (BeV) representations have emerged as the de-facto shared space in driving applications, offering a unified space for sensor data fusion and supporting various downstream tasks. However, conventional models use grids with fixed resolution and range and face computational inefficiencies due to the uniform allocation of resources across all cells. To address this, we propose PointBeV, a novel sparse BeV segmentation model operating on sparse BeV cells instead of dense grids. This approach offers precise control over memory usage, enabling the use of long temporal contexts and accommodating memory-constrained platforms. PointBeV employs an efficient two-pass strategy for training, enabling focused computation on regions of interest. At inference time, it can be used with various memory/performance trade-offs and flexibly adjusts to new specific use cases. PointBeV achieves state-of-the-art results on the nuScenes dataset for vehicle, pedestrian, and lane segmentation, showcasing superior performance in static and temporal settings despite being trained solely with sparse signals. We will release our code along with two new efficient modules used in the architecture: Sparse Feature Pulling, designed for the effective extraction of features from images to BeV, and Submanifold Attention, which enables efficient temporal modeling. Our code is available at https://github.com/valeoai/PointBeV.
title PointBeV: A Sparse Approach to BeV Predictions
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.00703