ReasonDrive: Efficient Visual Question Answering for Autonomous Vehicles with Reasoning-Enhanced Small Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Chahe, Amirhosein, Zhou, Lifeng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Query3D: LLM-Powered Open-Vocabulary Scene Segmentation with Language Embedded 3D Gaussian
di: Chahe, Amirhosein, et al.
Pubblicazione: (2024)
di: Chahe, Amirhosein, et al.
Pubblicazione: (2024)
Dynamic Adversarial Attacks on Autonomous Driving Systems
di: Chahe, Amirhosein, et al.
Pubblicazione: (2023)
di: Chahe, Amirhosein, et al.
Pubblicazione: (2023)
Policy-Guided World Model Planning for Language-Conditioned Visual Navigation
di: Chahe, Amirhosein, et al.
Pubblicazione: (2026)
di: Chahe, Amirhosein, et al.
Pubblicazione: (2026)
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
di: Xing, Shuo, et al.
Pubblicazione: (2024)
di: Xing, Shuo, et al.
Pubblicazione: (2024)
Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
di: Zhang, Dapeng, et al.
Pubblicazione: (2025)
di: Zhang, Dapeng, et al.
Pubblicazione: (2025)
DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes
di: Park, Yohan, et al.
Pubblicazione: (2025)
di: Park, Yohan, et al.
Pubblicazione: (2025)
VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous Driving
di: Zhang, Haiming, et al.
Pubblicazione: (2024)
di: Zhang, Haiming, et al.
Pubblicazione: (2024)
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
di: Gao, Haoxiang, et al.
Pubblicazione: (2025)
di: Gao, Haoxiang, et al.
Pubblicazione: (2025)
Vision and Language: Novel Representations and Artificial intelligence for Driving Scene Safety Assessment and Autonomous Vehicle Planning
di: Greer, Ross, et al.
Pubblicazione: (2026)
di: Greer, Ross, et al.
Pubblicazione: (2026)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
di: Zhao, Qingqing, et al.
Pubblicazione: (2025)
di: Zhao, Qingqing, et al.
Pubblicazione: (2025)
Language-Enhanced Latent Representations for Out-of-Distribution Detection in Autonomous Driving
di: Mao, Zhenjiang, et al.
Pubblicazione: (2024)
di: Mao, Zhenjiang, et al.
Pubblicazione: (2024)
ESPIRE: A Diagnostic Benchmark for Embodied Spatial Reasoning of Vision-Language Models
di: Zhao, Yanpeng, et al.
Pubblicazione: (2026)
di: Zhao, Yanpeng, et al.
Pubblicazione: (2026)
Machine Learning-Based Vehicle Intention Trajectory Recognition and Prediction for Autonomous Driving
di: Yu, Hanyi, et al.
Pubblicazione: (2024)
di: Yu, Hanyi, et al.
Pubblicazione: (2024)
LLM-attacker: Enhancing Closed-loop Adversarial Scenario Generation for Autonomous Driving with Large Language Models
di: Mei, Yuewen, et al.
Pubblicazione: (2025)
di: Mei, Yuewen, et al.
Pubblicazione: (2025)
Perception Without Vision for Trajectory Prediction: Ego Vehicle Dynamics as Scene Representation for Efficient Active Learning in Autonomous Driving
di: Greer, Ross, et al.
Pubblicazione: (2024)
di: Greer, Ross, et al.
Pubblicazione: (2024)
DeeAD: Dynamic Early Exit of Vision-Language Action for Efficient Autonomous Driving
di: HU, Haibo, et al.
Pubblicazione: (2025)
di: HU, Haibo, et al.
Pubblicazione: (2025)
LingoQA: Visual Question Answering for Autonomous Driving
di: Marcu, Ana-Maria, et al.
Pubblicazione: (2023)
di: Marcu, Ana-Maria, et al.
Pubblicazione: (2023)
Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving
di: Sun, Zhengqi, et al.
Pubblicazione: (2026)
di: Sun, Zhengqi, et al.
Pubblicazione: (2026)
DINO Pre-training for Vision-based End-to-end Autonomous Driving
di: Juneja, Shubham, et al.
Pubblicazione: (2024)
di: Juneja, Shubham, et al.
Pubblicazione: (2024)
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
di: Chen, Boyuan, et al.
Pubblicazione: (2024)
di: Chen, Boyuan, et al.
Pubblicazione: (2024)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
di: Goetting, Dylan, et al.
Pubblicazione: (2024)
di: Goetting, Dylan, et al.
Pubblicazione: (2024)
Automated Data Curation Using GPS & NLP to Generate Instruction-Action Pairs for Autonomous Vehicle Vision-Language Navigation Datasets
di: Roque, Guillermo, et al.
Pubblicazione: (2025)
di: Roque, Guillermo, et al.
Pubblicazione: (2025)
A Survey for Foundation Models in Autonomous Driving
di: Gao, Haoxiang, et al.
Pubblicazione: (2024)
di: Gao, Haoxiang, et al.
Pubblicazione: (2024)
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
di: Huang, Chi-Pin, et al.
Pubblicazione: (2025)
di: Huang, Chi-Pin, et al.
Pubblicazione: (2025)
OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving
di: Xing, Shuo, et al.
Pubblicazione: (2024)
di: Xing, Shuo, et al.
Pubblicazione: (2024)
Advancing Egocentric Video Question Answering with Multimodal Large Language Models
di: Patel, Alkesh, et al.
Pubblicazione: (2025)
di: Patel, Alkesh, et al.
Pubblicazione: (2025)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
di: Kong, Lingdong, et al.
Pubblicazione: (2024)
di: Kong, Lingdong, et al.
Pubblicazione: (2024)
PVI: Plug-in Visual Injection for Vision-Language-Action Models
di: Zhang, Zezhou, et al.
Pubblicazione: (2026)
di: Zhang, Zezhou, et al.
Pubblicazione: (2026)
Explore until Confident: Efficient Exploration for Embodied Question Answering
di: Ren, Allen Z., et al.
Pubblicazione: (2024)
di: Ren, Allen Z., et al.
Pubblicazione: (2024)
Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning
di: Huang, Chi-Pin, et al.
Pubblicazione: (2026)
di: Huang, Chi-Pin, et al.
Pubblicazione: (2026)
Addressing the Waypoint-Action Gap in End-to-End Autonomous Driving via Vehicle Motion Models
di: Rodríguez-Vidal, Jorge Daniel, et al.
Pubblicazione: (2026)
di: Rodríguez-Vidal, Jorge Daniel, et al.
Pubblicazione: (2026)
Test-Time Training for Visual Foresight Vision-Language-Action Models
di: Park, Sangwu, et al.
Pubblicazione: (2026)
di: Park, Sangwu, et al.
Pubblicazione: (2026)
DriveTransformer: Unified Transformer for Scalable End-to-End Autonomous Driving
di: Jia, Xiaosong, et al.
Pubblicazione: (2025)
di: Jia, Xiaosong, et al.
Pubblicazione: (2025)
ACT-Bench: Towards Action Controllable World Models for Autonomous Driving
di: Arai, Hidehisa, et al.
Pubblicazione: (2024)
di: Arai, Hidehisa, et al.
Pubblicazione: (2024)
Natural Language Instructions for Scene-Responsive Human-in-the-Loop Motion Planning in Autonomous Driving using Vision-Language-Action Models
di: Martinez-Sanchez, Angel, et al.
Pubblicazione: (2026)
di: Martinez-Sanchez, Angel, et al.
Pubblicazione: (2026)
Panoptic Perception for Autonomous Driving: A Survey
di: Li, Yunge, et al.
Pubblicazione: (2024)
di: Li, Yunge, et al.
Pubblicazione: (2024)
Hybrid-Prediction Integrated Planning for Autonomous Driving
di: Liu, Haochen, et al.
Pubblicazione: (2024)
di: Liu, Haochen, et al.
Pubblicazione: (2024)
RALACs: Action Recognition in Autonomous Vehicles using Interaction Encoding and Optical Flow
di: Zhou, Eddy, et al.
Pubblicazione: (2022)
di: Zhou, Eddy, et al.
Pubblicazione: (2022)
See Less, Drive Better: Generalizable End-to-End Autonomous Driving via Foundation Models Stochastic Patch Selection
di: Mallak, Amir, et al.
Pubblicazione: (2026)
di: Mallak, Amir, et al.
Pubblicazione: (2026)
Bench2Drive-R: Turning Real World Data into Reactive Closed-Loop Autonomous Driving Benchmark by Generative Model
di: You, Junqi, et al.
Pubblicazione: (2024)
di: You, Junqi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Query3D: LLM-Powered Open-Vocabulary Scene Segmentation with Language Embedded 3D Gaussian
di: Chahe, Amirhosein, et al.
Pubblicazione: (2024) -
Dynamic Adversarial Attacks on Autonomous Driving Systems
di: Chahe, Amirhosein, et al.
Pubblicazione: (2023) -
Policy-Guided World Model Planning for Language-Conditioned Visual Navigation
di: Chahe, Amirhosein, et al.
Pubblicazione: (2026) -
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
di: Xing, Shuo, et al.
Pubblicazione: (2024) -
Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
di: Zhang, Dapeng, et al.
Pubblicazione: (2025)