doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation
Fuente:
arXiv
Saved in:
| Main Authors: | Roy, Parthib, Perisetla, Srinivasa, Shriram, Shashank, Krishnaswamy, Harsha, Keskar, Aryan, Greer, Ross |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety
by: Shriram, Shashank, et al.
Published: (2025)
by: Shriram, Shashank, et al.
Published: (2025)
Vision and Language: Novel Representations and Artificial intelligence for Driving Scene Safety Assessment and Autonomous Vehicle Planning
by: Greer, Ross, et al.
Published: (2026)
by: Greer, Ross, et al.
Published: (2026)
Natural Language Instructions for Scene-Responsive Human-in-the-Loop Motion Planning in Autonomous Driving using Vision-Language-Action Models
by: Martinez-Sanchez, Angel, et al.
Published: (2026)
by: Martinez-Sanchez, Angel, et al.
Published: (2026)
Automated Data Curation Using GPS & NLP to Generate Instruction-Action Pairs for Autonomous Vehicle Vision-Language Navigation Datasets
by: Roque, Guillermo, et al.
Published: (2025)
by: Roque, Guillermo, et al.
Published: (2025)
Perception Without Vision for Trajectory Prediction: Ego Vehicle Dynamics as Scene Representation for Efficient Active Learning in Autonomous Driving
by: Greer, Ross, et al.
Published: (2024)
by: Greer, Ross, et al.
Published: (2024)
Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving
by: Gopalkrishnan, Akshay, et al.
Published: (2024)
by: Gopalkrishnan, Akshay, et al.
Published: (2024)
MTR-VP: Towards End-to-End Trajectory Planning through Context-Driven Image Encoding and Multiple Trajectory Prediction
by: Keskar, Maitrayee, et al.
Published: (2025)
by: Keskar, Maitrayee, et al.
Published: (2025)
DepthVision: Enabling Robust Vision-Language Models with GAN-Based LiDAR-to-RGB Synthesis for Autonomous Driving
by: Kirchner, Sven, et al.
Published: (2025)
by: Kirchner, Sven, et al.
Published: (2025)
Towards Explainable, Safe Autonomous Driving with Language Embeddings for Novelty Identification and Active Learning: Framework and Experimental Analysis with Real-World Data Sets
by: Greer, Ross, et al.
Published: (2024)
by: Greer, Ross, et al.
Published: (2024)
Vision-Based Natural Language Scene Understanding for Autonomous Driving: An Extended Dataset and a New Model for Traffic Scene Description Generation
by: Zadeh, Danial Sadrian, et al.
Published: (2026)
by: Zadeh, Danial Sadrian, et al.
Published: (2026)
Can Vision-Language Models Understand and Interpret Dynamic Gestures from Pedestrians? Pilot Datasets and Exploration Towards Instructive Nonverbal Commands for Cooperative Autonomous Vehicles
by: Bossen, Tonko E. W., et al.
Published: (2025)
by: Bossen, Tonko E. W., et al.
Published: (2025)
Evaluating Vision-Language Models for Zero-Shot Detection, Classification, and Association of Motorcycles, Passengers, and Helmets
by: Choi, Lucas, et al.
Published: (2024)
by: Choi, Lucas, et al.
Published: (2024)
Evaluating Cascaded Methods of Vision-Language Models for Zero-Shot Detection and Association of Hardhats for Increased Construction Safety
by: Choi, Lucas, et al.
Published: (2024)
by: Choi, Lucas, et al.
Published: (2024)
Beyond General Prompts: Automated Prompt Refinement using Contrastive Class Alignment Scores for Disambiguating Objects in Vision-Language Models
by: Choi, Lucas, et al.
Published: (2025)
by: Choi, Lucas, et al.
Published: (2025)
Looking and Listening Inside and Outside: Multimodal Artificial Intelligence Systems for Driver Safety Assessment and Intelligent Vehicle Decision-Making
by: Greer, Ross, et al.
Published: (2026)
by: Greer, Ross, et al.
Published: (2026)
AirNav: A Large-Scale UAV Vision-and-Language Navigation Dataset with Natural and Diverse Instructions
by: Cai, Hengxing, et al.
Published: (2026)
by: Cai, Hengxing, et al.
Published: (2026)
Grounded Concreteness: Human-Like Concreteness Sensitivity in Vision-Language Models
by: Roy, Aryan, et al.
Published: (2026)
by: Roy, Aryan, et al.
Published: (2026)
Words to Wheels: Vision-Based Autonomous Driving Understanding Human Language Instructions Using Foundation Models
by: Ryu, Chanhoe, et al.
Published: (2024)
by: Ryu, Chanhoe, et al.
Published: (2024)
Application of Vision-Language Model to Pedestrians Behavior and Scene Understanding in Autonomous Driving
by: Gao, Haoxiang, et al.
Published: (2025)
by: Gao, Haoxiang, et al.
Published: (2025)
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding
by: Elhenawy, Mohammed, et al.
Published: (2025)
by: Elhenawy, Mohammed, et al.
Published: (2025)
RAD-LAD: Rule and Language Grounded Autonomous Driving in Real-Time
by: Ghosh, Anurag, et al.
Published: (2026)
by: Ghosh, Anurag, et al.
Published: (2026)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
by: Wang, Shihao, et al.
Published: (2024)
by: Wang, Shihao, et al.
Published: (2024)
CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving
by: Arai, Hidehisa, et al.
Published: (2024)
by: Arai, Hidehisa, et al.
Published: (2024)
Agro-Consensus: Semantic Self-Consistency in Vision-Language Models for Crop Disease Management in Developing Countries
by: Gupta, Mihir, et al.
Published: (2025)
by: Gupta, Mihir, et al.
Published: (2025)
Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
Situation-Aware Feedback-Predictive Control Framework for Lane-Less Dense Traffic
by: Khound, Parthib
Published: (2026)
by: Khound, Parthib
Published: (2026)
VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
ScenePilot-4K: A Large-Scale First-Person Dataset and Benchmark for Vision-Language Models in Autonomous Driving
by: Wang, Yujin, et al.
Published: (2026)
by: Wang, Yujin, et al.
Published: (2026)
General Scene Adaptation for Vision-and-Language Navigation
by: Hong, Haodong, et al.
Published: (2025)
by: Hong, Haodong, et al.
Published: (2025)
Vega: Learning to Drive with Natural Language Instructions
by: Zuo, Sicheng, et al.
Published: (2026)
by: Zuo, Sicheng, et al.
Published: (2026)
Cross-Lingual Transfer Robustness to Lower-Resource Languages on Adversarial Datasets
by: Manafi, Shadi, et al.
Published: (2024)
by: Manafi, Shadi, et al.
Published: (2024)
Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed Environments
by: Hong, Haodong, et al.
Published: (2024)
by: Hong, Haodong, et al.
Published: (2024)
Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
by: Zhao, Zongchuang, et al.
Published: (2025)
by: Zhao, Zongchuang, et al.
Published: (2025)
Driver Activity Classification Using Generalizable Representations from Vision-Language Models
by: Greer, Ross, et al.
Published: (2024)
by: Greer, Ross, et al.
Published: (2024)
Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving
by: Hu, Haibo, et al.
Published: (2025)
by: Hu, Haibo, et al.
Published: (2025)
Profiling Programming Language Learning
by: Crichton, Will, et al.
Published: (2024)
by: Crichton, Will, et al.
Published: (2024)
Self-Consistency in Vision-Language Models for Precision Agriculture: Multi-Response Consensus for Crop Disease Management
by: Gupta, Mihir, et al.
Published: (2025)
by: Gupta, Mihir, et al.
Published: (2025)
VLP: Vision Language Planning for Autonomous Driving
by: Pan, Chenbin, et al.
Published: (2024)
by: Pan, Chenbin, et al.
Published: (2024)
ESceme: Vision-and-Language Navigation with Episodic Scene Memory
by: Zheng, Qi, et al.
Published: (2023)
by: Zheng, Qi, et al.
Published: (2023)
Similar Items
-
Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety
by: Shriram, Shashank, et al.
Published: (2025) -
Vision and Language: Novel Representations and Artificial intelligence for Driving Scene Safety Assessment and Autonomous Vehicle Planning
by: Greer, Ross, et al.
Published: (2026) -
Natural Language Instructions for Scene-Responsive Human-in-the-Loop Motion Planning in Autonomous Driving using Vision-Language-Action Models
by: Martinez-Sanchez, Angel, et al.
Published: (2026) -
Automated Data Curation Using GPS & NLP to Generate Instruction-Action Pairs for Autonomous Vehicle Vision-Language Navigation Datasets
by: Roque, Guillermo, et al.
Published: (2025) -
Perception Without Vision for Trajectory Prediction: Ego Vehicle Dynamics as Scene Representation for Efficient Active Learning in Autonomous Driving
by: Greer, Ross, et al.
Published: (2024)