Reconstruction Matters: Learning Geometry-Aligned BEV Representation through 3D Gaussian Splatting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lu, Yiren, Ye, Xin, Yaman, Burhaneddin, Luo, Jingru, Xiong, Zhexiao, Ren, Liu, Yin, Yu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917353672933376
author Lu, Yiren
Ye, Xin
Yaman, Burhaneddin
Luo, Jingru
Xiong, Zhexiao
Ren, Liu
Yin, Yu
author_facet Lu, Yiren
Ye, Xin
Yaman, Burhaneddin
Luo, Jingru
Xiong, Zhexiao
Ren, Liu
Yin, Yu
contents Bird's-Eye-View (BEV) perception serves as a cornerstone for autonomous driving, offering a unified spatial representation that fuses surrounding-view images to enable reasoning for various downstream tasks, such as semantic segmentation, 3D object detection, and motion prediction. However, most existing BEV perception frameworks adopt an end-to-end training paradigm, where image features are directly transformed into the BEV space and optimized solely through downstream task supervision. This formulation treats the entire perception process as a black box, often lacking explicit 3D geometric understanding and interpretability, leading to suboptimal performance. In this paper, we claim that an explicit 3D representation matters for accurate BEV perception, and we propose Splat2BEV, a Gaussian Splatting-assisted framework for BEV tasks. Splat2BEV aims to learn BEV feature representations that are both semantically rich and geometrically precise. We first pre-train a Gaussian generator that explicitly reconstructs 3D scenes from multi-view inputs, enabling the generation of geometry-aligned feature representations. These representations are then projected into the BEV space to serve as inputs for downstream tasks. Extensive experiments on nuScenes and argoverse dataset demonstrate that Splat2BEV achieves state-of-the-art performance and validate the effectiveness of incorporating explicit 3D reconstruction into BEV perception.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19193
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Reconstruction Matters: Learning Geometry-Aligned BEV Representation through 3D Gaussian Splatting
Lu, Yiren
Ye, Xin
Yaman, Burhaneddin
Luo, Jingru
Xiong, Zhexiao
Ren, Liu
Yin, Yu
Computer Vision and Pattern Recognition
Bird's-Eye-View (BEV) perception serves as a cornerstone for autonomous driving, offering a unified spatial representation that fuses surrounding-view images to enable reasoning for various downstream tasks, such as semantic segmentation, 3D object detection, and motion prediction. However, most existing BEV perception frameworks adopt an end-to-end training paradigm, where image features are directly transformed into the BEV space and optimized solely through downstream task supervision. This formulation treats the entire perception process as a black box, often lacking explicit 3D geometric understanding and interpretability, leading to suboptimal performance. In this paper, we claim that an explicit 3D representation matters for accurate BEV perception, and we propose Splat2BEV, a Gaussian Splatting-assisted framework for BEV tasks. Splat2BEV aims to learn BEV feature representations that are both semantically rich and geometrically precise. We first pre-train a Gaussian generator that explicitly reconstructs 3D scenes from multi-view inputs, enabling the generation of geometry-aligned feature representations. These representations are then projected into the BEV space to serve as inputs for downstream tasks. Extensive experiments on nuScenes and argoverse dataset demonstrate that Splat2BEV achieves state-of-the-art performance and validate the effectiveness of incorporating explicit 3D reconstruction into BEV perception.
title Reconstruction Matters: Learning Geometry-Aligned BEV Representation through 3D Gaussian Splatting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.19193