An Initial Study of Bird's-Eye View Generation for Autonomous Vehicles using Cross-View Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916904828928000 |
|---|---|
| author | Santos, Felipe Carlos dos Antonelo, Eric Aislan Couto, Gustavo Claudio Karl |
| author_facet | Santos, Felipe Carlos dos Antonelo, Eric Aislan Couto, Gustavo Claudio Karl |
| contents | Bird's-Eye View (BEV) maps provide a structured, top-down abstraction that is crucial for autonomous-driving perception. In this work, we employ Cross-View Transformers (CVT) for learning to map camera images to three BEV's channels - road, lane markings, and planned trajectory - using a realistic simulator for urban driving. Our study examines generalization to unseen towns, the effect of different camera layouts, and two loss formulations (focal and L1). Using training data from only a town, a four-camera CVT trained with the L1 loss delivers the most robust test performance, evaluated in a new town. Overall, our results underscore CVT's promise for mapping camera inputs to reasonably accurate BEV maps. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_12520 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | An Initial Study of Bird's-Eye View Generation for Autonomous Vehicles using Cross-View Transformers Santos, Felipe Carlos dos Antonelo, Eric Aislan Couto, Gustavo Claudio Karl Computer Vision and Pattern Recognition Artificial Intelligence Bird's-Eye View (BEV) maps provide a structured, top-down abstraction that is crucial for autonomous-driving perception. In this work, we employ Cross-View Transformers (CVT) for learning to map camera images to three BEV's channels - road, lane markings, and planned trajectory - using a realistic simulator for urban driving. Our study examines generalization to unseen towns, the effect of different camera layouts, and two loss formulations (focal and L1). Using training data from only a town, a four-camera CVT trained with the L1 loss delivers the most robust test performance, evaluated in a new town. Overall, our results underscore CVT's promise for mapping camera inputs to reasonably accurate BEV maps. |
| title | An Initial Study of Bird's-Eye View Generation for Autonomous Vehicles using Cross-View Transformers |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2508.12520 |