BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Monninger, Thomas, Xie, Shaoyuan, Chen, Qi Alfred, Ding, Sihao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NavMapFusion: Diffusion-based Fusion of Navigation Maps for Online Vectorized HD Map Construction
von: Monninger, Thomas, et al.
Veröffentlicht: (2025)
von: Monninger, Thomas, et al.
Veröffentlicht: (2025)
AugMapNet: Improving Spatial Latent Structure via BEV Grid Augmentation for Enhanced Vectorized Online HD Map Construction
von: Monninger, Thomas, et al.
Veröffentlicht: (2025)
von: Monninger, Thomas, et al.
Veröffentlicht: (2025)
MapDiffusion: Generative Diffusion for Vectorized Online HD Map Construction and Uncertainty Estimation in Autonomous Driving
von: Monninger, Thomas, et al.
Veröffentlicht: (2025)
von: Monninger, Thomas, et al.
Veröffentlicht: (2025)
LMT-Net: Lane Model Transformer Network for Automated HD Mapping from Sparse Vehicle Observations
von: Mink, Michael, et al.
Veröffentlicht: (2024)
von: Mink, Michael, et al.
Veröffentlicht: (2024)
AutoVDC: Automated Vision Data Cleaning Using Vision-Language Models
von: Vasa, Santosh, et al.
Veröffentlicht: (2025)
von: Vasa, Santosh, et al.
Veröffentlicht: (2025)
TempBEV: Improving Learned BEV Encoders with Combined Image and BEV Space Temporal Aggregation
von: Monninger, Thomas, et al.
Veröffentlicht: (2024)
von: Monninger, Thomas, et al.
Veröffentlicht: (2024)
Benchmarking and Improving Bird's Eye View Perception Robustness in Autonomous Driving
von: Xie, Shaoyuan, et al.
Veröffentlicht: (2024)
von: Xie, Shaoyuan, et al.
Veröffentlicht: (2024)
nuCarla: A nuScenes-Style Bird's-Eye View Perception Dataset for CARLA Simulation
von: Qiao, Zhijie, et al.
Veröffentlicht: (2025)
von: Qiao, Zhijie, et al.
Veröffentlicht: (2025)
Bird Eye-View to Street-View: A Survey
von: Bajbaa, Khawlah, et al.
Veröffentlicht: (2024)
von: Bajbaa, Khawlah, et al.
Veröffentlicht: (2024)
Distilling Knowledge for Short-to-Long Term Trajectory Prediction
von: Das, Sourav, et al.
Veröffentlicht: (2023)
von: Das, Sourav, et al.
Veröffentlicht: (2023)
Improving Bird's Eye View Semantic Segmentation by Task Decomposition
von: Zhao, Tianhao, et al.
Veröffentlicht: (2024)
von: Zhao, Tianhao, et al.
Veröffentlicht: (2024)
KDMOS:Knowledge Distillation for Motion Segmentation
von: Cao, Chunyu, et al.
Veröffentlicht: (2025)
von: Cao, Chunyu, et al.
Veröffentlicht: (2025)
LIX: Implicitly Infusing Spatial Geometric Prior Knowledge into Visual Semantic Segmentation for Autonomous Driving
von: Guo, Sicen, et al.
Veröffentlicht: (2024)
von: Guo, Sicen, et al.
Veröffentlicht: (2024)
Epipolar Attention Field Transformers for Bird's Eye View Semantic Segmentation
von: Witte, Christian, et al.
Veröffentlicht: (2024)
von: Witte, Christian, et al.
Veröffentlicht: (2024)
M2Distill: Multi-Modal Distillation for Lifelong Imitation Learning
von: Roy, Kaushik, et al.
Veröffentlicht: (2024)
von: Roy, Kaushik, et al.
Veröffentlicht: (2024)
SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy Prediction
von: Duan, Zaipeng, et al.
Veröffentlicht: (2025)
von: Duan, Zaipeng, et al.
Veröffentlicht: (2025)
Diffusion Meets DAgger: Supercharging Eye-in-hand Imitation Learning
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2024)
Extrapolated Urban View Synthesis Benchmark
von: Han, Xiangyu, et al.
Veröffentlicht: (2024)
von: Han, Xiangyu, et al.
Veröffentlicht: (2024)
View-Invariant Policy Learning via Zero-Shot Novel View Synthesis
von: Tian, Stephen, et al.
Veröffentlicht: (2024)
von: Tian, Stephen, et al.
Veröffentlicht: (2024)
CycleBEV: Regularizing View Transformation Networks via View Cycle Consistency for Bird's-Eye-View Semantic Segmentation
von: Hong, Jeongbin, et al.
Veröffentlicht: (2026)
von: Hong, Jeongbin, et al.
Veröffentlicht: (2026)
Semantically Controllable Augmentations for Generalizable Robot Learning
von: Chen, Zoey, et al.
Veröffentlicht: (2024)
von: Chen, Zoey, et al.
Veröffentlicht: (2024)
PolarBEVDet: Exploring Polar Representation for Multi-View 3D Object Detection in Bird's-Eye-View
von: Yu, Zichen, et al.
Veröffentlicht: (2024)
von: Yu, Zichen, et al.
Veröffentlicht: (2024)
Theia: Distilling Diverse Vision Foundation Models for Robot Learning
von: Shang, Jinghuan, et al.
Veröffentlicht: (2024)
von: Shang, Jinghuan, et al.
Veröffentlicht: (2024)
LetsMap: Unsupervised Representation Learning for Semantic BEV Mapping
von: Gosala, Nikhil, et al.
Veröffentlicht: (2024)
von: Gosala, Nikhil, et al.
Veröffentlicht: (2024)
BEVCALIB: LiDAR-Camera Calibration via Geometry-Guided Bird's-Eye View Representations
von: Yuan, Weiduo, et al.
Veröffentlicht: (2025)
von: Yuan, Weiduo, et al.
Veröffentlicht: (2025)
Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery
von: Tran, Chi-Nguyen, et al.
Veröffentlicht: (2026)
von: Tran, Chi-Nguyen, et al.
Veröffentlicht: (2026)
DuoSpaceNet: Leveraging Both Bird's-Eye-View and Perspective View Representations for 3D Object Detection
von: Huang, Zhe, et al.
Veröffentlicht: (2024)
von: Huang, Zhe, et al.
Veröffentlicht: (2024)
Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis
von: Van Hoorick, Basile, et al.
Veröffentlicht: (2024)
von: Van Hoorick, Basile, et al.
Veröffentlicht: (2024)
ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
von: Liu, Shaowei, et al.
Veröffentlicht: (2025)
von: Liu, Shaowei, et al.
Veröffentlicht: (2025)
Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent
von: He, Linfeng, et al.
Veröffentlicht: (2024)
von: He, Linfeng, et al.
Veröffentlicht: (2024)
MVSA-Net: Multi-View State-Action Recognition for Robust and Deployable Trajectory Generation
von: Asali, Ehsan, et al.
Veröffentlicht: (2023)
von: Asali, Ehsan, et al.
Veröffentlicht: (2023)
Hyperspectral Adapter for Semantic Segmentation with Vision Foundation Models
von: Hurtado, Juana Valeria, et al.
Veröffentlicht: (2025)
von: Hurtado, Juana Valeria, et al.
Veröffentlicht: (2025)
How to Benchmark Vision Foundation Models for Semantic Segmentation?
von: Kerssies, Tommie, et al.
Veröffentlicht: (2024)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2024)
ForecastOcc: Vision-based Semantic Occupancy Forecasting
von: Mohan, Riya, et al.
Veröffentlicht: (2026)
von: Mohan, Riya, et al.
Veröffentlicht: (2026)
Visual Representation Learning with Stochastic Frame Prediction
von: Jang, Huiwon, et al.
Veröffentlicht: (2024)
von: Jang, Huiwon, et al.
Veröffentlicht: (2024)
CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining
von: Liu, I-Chun Arthur, et al.
Veröffentlicht: (2026)
von: Liu, I-Chun Arthur, et al.
Veröffentlicht: (2026)
Learning Content-Aware Multi-Modal Joint Input Pruning via Bird's-Eye-View Representation
von: Li, Yuxin, et al.
Veröffentlicht: (2024)
von: Li, Yuxin, et al.
Veröffentlicht: (2024)
SplaTraj: Camera Trajectory Generation with Semantic Gaussian Splatting
von: Liu, Xinyi, et al.
Veröffentlicht: (2024)
von: Liu, Xinyi, et al.
Veröffentlicht: (2024)
EvObj: Learning Evolving Object-centric Representations for 3D Instance Segmentation without Scene Supervision
von: Chen, Jiahao, et al.
Veröffentlicht: (2026)
von: Chen, Jiahao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
NavMapFusion: Diffusion-based Fusion of Navigation Maps for Online Vectorized HD Map Construction
von: Monninger, Thomas, et al.
Veröffentlicht: (2025) -
AugMapNet: Improving Spatial Latent Structure via BEV Grid Augmentation for Enhanced Vectorized Online HD Map Construction
von: Monninger, Thomas, et al.
Veröffentlicht: (2025) -
MapDiffusion: Generative Diffusion for Vectorized Online HD Map Construction and Uncertainty Estimation in Autonomous Driving
von: Monninger, Thomas, et al.
Veröffentlicht: (2025) -
LMT-Net: Lane Model Transformer Network for Automated HD Mapping from Sparse Vehicle Observations
von: Mink, Michael, et al.
Veröffentlicht: (2024) -
AutoVDC: Automated Vision Data Cleaning Using Vision-Language Models
von: Vasa, Santosh, et al.
Veröffentlicht: (2025)