An Initial Study of Bird's-Eye View Generation for Autonomous Vehicles using Cross-View Transformers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Santos, Felipe Carlos dos, Antonelo, Eric Aislan, Couto, Gustavo Claudio Karl
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916904828928000
author Santos, Felipe Carlos dos
Antonelo, Eric Aislan
Couto, Gustavo Claudio Karl
author_facet Santos, Felipe Carlos dos
Antonelo, Eric Aislan
Couto, Gustavo Claudio Karl
contents Bird's-Eye View (BEV) maps provide a structured, top-down abstraction that is crucial for autonomous-driving perception. In this work, we employ Cross-View Transformers (CVT) for learning to map camera images to three BEV's channels - road, lane markings, and planned trajectory - using a realistic simulator for urban driving. Our study examines generalization to unseen towns, the effect of different camera layouts, and two loss formulations (focal and L1). Using training data from only a town, a four-camera CVT trained with the L1 loss delivers the most robust test performance, evaluated in a new town. Overall, our results underscore CVT's promise for mapping camera inputs to reasonably accurate BEV maps.
format Preprint
id arxiv_https___arxiv_org_abs_2508_12520
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Initial Study of Bird's-Eye View Generation for Autonomous Vehicles using Cross-View Transformers
Santos, Felipe Carlos dos
Antonelo, Eric Aislan
Couto, Gustavo Claudio Karl
Computer Vision and Pattern Recognition
Artificial Intelligence
Bird's-Eye View (BEV) maps provide a structured, top-down abstraction that is crucial for autonomous-driving perception. In this work, we employ Cross-View Transformers (CVT) for learning to map camera images to three BEV's channels - road, lane markings, and planned trajectory - using a realistic simulator for urban driving. Our study examines generalization to unseen towns, the effect of different camera layouts, and two loss formulations (focal and L1). Using training data from only a town, a four-camera CVT trained with the L1 loss delivers the most robust test performance, evaluated in a new town. Overall, our results underscore CVT's promise for mapping camera inputs to reasonably accurate BEV maps.
title An Initial Study of Bird's-Eye View Generation for Autonomous Vehicles using Cross-View Transformers
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2508.12520