Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent
Fuente:
arXiv
Saved in:
| Main Authors: | He, Linfeng, Sun, Yiming, Wu, Sihao, Liu, Jiaxu, Huang, Xiaowei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Autonomous Driving Safety with Collision Scenario Integration
by: Wang, Zi, et al.
Published: (2025)
by: Wang, Zi, et al.
Published: (2025)
Asynchronous Large Language Model Enhanced Planner for Autonomous Driving
by: Chen, Yuan, et al.
Published: (2024)
by: Chen, Yuan, et al.
Published: (2024)
Scalable Vision-Based 3D Object Detection and Monocular Depth Estimation for Autonomous Driving
by: Liu, Yuxuan
Published: (2024)
by: Liu, Yuxuan
Published: (2024)
Data Shift of Object Detection in Autonomous Driving
by: Xu, Lida
Published: (2025)
by: Xu, Lida
Published: (2025)
Hybrid-Prediction Integrated Planning for Autonomous Driving
by: Liu, Haochen, et al.
Published: (2024)
by: Liu, Haochen, et al.
Published: (2024)
Cross-Cluster Shifting for Efficient and Effective 3D Object Detection in Autonomous Driving
by: Chen, Zhili, et al.
Published: (2024)
by: Chen, Zhili, et al.
Published: (2024)
Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
by: Jiang, Bo, et al.
Published: (2024)
by: Jiang, Bo, et al.
Published: (2024)
LOID: Lane Occlusion Inpainting and Detection for Enhanced Autonomous Driving Systems
by: Agrawal, Aayush, et al.
Published: (2024)
by: Agrawal, Aayush, et al.
Published: (2024)
Graph-Based Multi-Modal Sensor Fusion for Autonomous Driving
by: Sani, Depanshu, et al.
Published: (2024)
by: Sani, Depanshu, et al.
Published: (2024)
Language-Enhanced Latent Representations for Out-of-Distribution Detection in Autonomous Driving
by: Mao, Zhenjiang, et al.
Published: (2024)
by: Mao, Zhenjiang, et al.
Published: (2024)
3D Object Visibility Prediction in Autonomous Driving
by: Luo, Chuanyu, et al.
Published: (2024)
by: Luo, Chuanyu, et al.
Published: (2024)
Distilling Multi-modal Large Language Models for Autonomous Driving
by: Hegde, Deepti, et al.
Published: (2025)
by: Hegde, Deepti, et al.
Published: (2025)
OLiDM: Object-aware LiDAR Diffusion Models for Autonomous Driving
by: Yan, Tianyi, et al.
Published: (2024)
by: Yan, Tianyi, et al.
Published: (2024)
Enhanced Safety in Autonomous Driving: Integrating Latent State Diffusion Model for End-to-End Navigation
by: Chu, Detian, et al.
Published: (2024)
by: Chu, Detian, et al.
Published: (2024)
123D: Unifying Multi-Modal Autonomous Driving Data at Scale
by: Dauner, Daniel, et al.
Published: (2026)
by: Dauner, Daniel, et al.
Published: (2026)
MSSF: A 4D Radar and Camera Fusion Framework With Multi-Stage Sampling for 3D Object Detection in Autonomous Driving
by: Liu, Hongsi, et al.
Published: (2024)
by: Liu, Hongsi, et al.
Published: (2024)
AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
by: Qian, Kangan, et al.
Published: (2025)
by: Qian, Kangan, et al.
Published: (2025)
DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
by: jia, Feiyang, et al.
Published: (2026)
by: jia, Feiyang, et al.
Published: (2026)
Unraveling the Effects of Synthetic Data on End-to-End Autonomous Driving
by: Ge, Junhao, et al.
Published: (2025)
by: Ge, Junhao, et al.
Published: (2025)
Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation
by: Chen, Bolei, et al.
Published: (2025)
by: Chen, Bolei, et al.
Published: (2025)
Unifying Language-Action Understanding and Generation for Autonomous Driving
by: Wang, Xinyang, et al.
Published: (2026)
by: Wang, Xinyang, et al.
Published: (2026)
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
by: Wu, Yanhao, et al.
Published: (2026)
by: Wu, Yanhao, et al.
Published: (2026)
ReasonDrive: Efficient Visual Question Answering for Autonomous Vehicles with Reasoning-Enhanced Small Vision-Language Models
by: Chahe, Amirhosein, et al.
Published: (2025)
by: Chahe, Amirhosein, et al.
Published: (2025)
Exploring the Causality of End-to-End Autonomous Driving
by: Li, Jiankun, et al.
Published: (2024)
by: Li, Jiankun, et al.
Published: (2024)
FisheyeDetNet: 360° Surround view Fisheye Camera based Object Detection System for Autonomous Driving
by: Sistu, Ganesh, et al.
Published: (2024)
by: Sistu, Ganesh, et al.
Published: (2024)
DiffAD: A Unified Diffusion Modeling Approach for Autonomous Driving
by: Wang, Tao, et al.
Published: (2025)
by: Wang, Tao, et al.
Published: (2025)
DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model
by: Xu, Zhenhua, et al.
Published: (2023)
by: Xu, Zhenhua, et al.
Published: (2023)
U-ViLAR: Uncertainty-Aware Visual Localization for Autonomous Driving via Differentiable Association and Registration
by: Li, Xiaofan, et al.
Published: (2025)
by: Li, Xiaofan, et al.
Published: (2025)
LLM-attacker: Enhancing Closed-loop Adversarial Scenario Generation for Autonomous Driving with Large Language Models
by: Mei, Yuewen, et al.
Published: (2025)
by: Mei, Yuewen, et al.
Published: (2025)
Multi-Object Tracking with Camera-LiDAR Fusion for Autonomous Driving
by: Pieroni, Riccardo, et al.
Published: (2024)
by: Pieroni, Riccardo, et al.
Published: (2024)
Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
by: Han, Jianhua, et al.
Published: (2025)
by: Han, Jianhua, et al.
Published: (2025)
MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving
by: Yasarla, Rajeev, et al.
Published: (2026)
by: Yasarla, Rajeev, et al.
Published: (2026)
Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving
by: Liu, Qiqi, et al.
Published: (2026)
by: Liu, Qiqi, et al.
Published: (2026)
DynFlowDrive: Flow-Based Dynamic World Modeling for Autonomous Driving
by: Liu, Xiaolu, et al.
Published: (2026)
by: Liu, Xiaolu, et al.
Published: (2026)
A Language Agent for Autonomous Driving
by: Mao, Jiageng, et al.
Published: (2023)
by: Mao, Jiageng, et al.
Published: (2023)
BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object Detection
by: Zhang, Guowen, et al.
Published: (2025)
by: Zhang, Guowen, et al.
Published: (2025)
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
HawkDrive: A Transformer-driven Visual Perception System for Autonomous Driving in Night Scene
by: Guo, Ziang, et al.
Published: (2024)
by: Guo, Ziang, et al.
Published: (2024)
A Survey on Vision-Language-Action Models for Autonomous Driving
by: Jiang, Sicong, et al.
Published: (2025)
by: Jiang, Sicong, et al.
Published: (2025)
CoReVLA: A Dual-Stage End-to-End Autonomous Driving Framework for Long-Tail Scenarios via Collect-and-Refine
by: Fang, Shiyu, et al.
Published: (2025)
by: Fang, Shiyu, et al.
Published: (2025)
Similar Items
-
Enhancing Autonomous Driving Safety with Collision Scenario Integration
by: Wang, Zi, et al.
Published: (2025) -
Asynchronous Large Language Model Enhanced Planner for Autonomous Driving
by: Chen, Yuan, et al.
Published: (2024) -
Scalable Vision-Based 3D Object Detection and Monocular Depth Estimation for Autonomous Driving
by: Liu, Yuxuan
Published: (2024) -
Data Shift of Object Detection in Autonomous Driving
by: Xu, Lida
Published: (2025) -
Hybrid-Prediction Integrated Planning for Autonomous Driving
by: Liu, Haochen, et al.
Published: (2024)