ScenePilot-4K: A Large-Scale First-Person Dataset and Benchmark for Vision-Language Models in Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yujin, Zheng, Yutong, Fan, Wenxian, Wang, Tianyi, Chu, Hongqing, Zhang, Li, Gao, Bingzhao, Tian, Daxin, Wang, Jianqiang, Chen, Hong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RAC3: Retrieval-Augmented Corner Case Comprehension for Autonomous Driving with Vision-Language Models
by: Wang, Yujin, et al.
Published: (2024)
by: Wang, Yujin, et al.
Published: (2024)
RAD: Retrieval-Augmented Decision-Making of Meta-Actions with Vision-Language Models in Autonomous Driving
by: Wang, Yujin, et al.
Published: (2025)
by: Wang, Yujin, et al.
Published: (2025)
ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driving
by: Ruan, Qiyu, et al.
Published: (2026)
by: Ruan, Qiyu, et al.
Published: (2026)
KEPT: Knowledge-Enhanced Prediction of Trajectories from Consecutive Driving Frames with Vision-Language Models
by: Wang, Yujin, et al.
Published: (2025)
by: Wang, Yujin, et al.
Published: (2025)
BIDA: A Bi-level Interaction Decision-making Algorithm for Autonomous Vehicles in Dynamic Traffic Scenarios
by: Yu, Liyang, et al.
Published: (2025)
by: Yu, Liyang, et al.
Published: (2025)
Hallucination Elimination and Semantic Enhancement Framework for Vision-Language Models in Traffic Scenarios
by: Fan, Jiaqi, et al.
Published: (2024)
by: Fan, Jiaqi, et al.
Published: (2024)
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios
by: Fan, Jiaqi, et al.
Published: (2024)
by: Fan, Jiaqi, et al.
Published: (2024)
DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving
by: Diao, Muxi, et al.
Published: (2025)
by: Diao, Muxi, et al.
Published: (2025)
SceneFake: An Initial Dataset and Benchmarks for Scene Fake Audio Detection
by: Yi, Jiangyan, et al.
Published: (2022)
by: Yi, Jiangyan, et al.
Published: (2022)
SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving
by: Li, Yiming, et al.
Published: (2023)
by: Li, Yiming, et al.
Published: (2023)
DriveE2E: Closed-Loop Benchmark for End-to-End Autonomous Driving through Real-to-Simulation
by: Yu, Haibao, et al.
Published: (2025)
by: Yu, Haibao, et al.
Published: (2025)
AdaptiveAE: An Adaptive Exposure Strategy for HDR Capturing in Dynamic Scenes
by: Xu, Tianyi, et al.
Published: (2025)
by: Xu, Tianyi, et al.
Published: (2025)
PreGSU-A Generalized Traffic Scene Understanding Model for Autonomous Driving based on Pre-trained Graph Attention Network
by: Wang, Yuning, et al.
Published: (2024)
by: Wang, Yuning, et al.
Published: (2024)
RSUD20K: A Dataset for Road Scene Understanding In Autonomous Driving
by: Zunair, Hasib, et al.
Published: (2024)
by: Zunair, Hasib, et al.
Published: (2024)
Prospective Role of Foundation Models in Advancing Autonomous Vehicles
by: Wu, Jianhua, et al.
Published: (2023)
by: Wu, Jianhua, et al.
Published: (2023)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
by: Tian, Kexin, et al.
Published: (2025)
by: Tian, Kexin, et al.
Published: (2025)
WayveScenes101: A Dataset and Benchmark for Novel View Synthesis in Autonomous Driving
by: Zürn, Jannik, et al.
Published: (2024)
by: Zürn, Jannik, et al.
Published: (2024)
Learning A Unified Risk Map for Autonomous Driving in Partially Observable Environments
by: Jia, Jie, et al.
Published: (2026)
by: Jia, Jie, et al.
Published: (2026)
Advancing Off-Road Autonomous Driving: The Large-Scale ORAD-3D Dataset and Comprehensive Benchmarks
by: Min, Chen, et al.
Published: (2025)
by: Min, Chen, et al.
Published: (2025)
PADriver: Towards Personalized Autonomous Driving
by: Kou, Genghua, et al.
Published: (2025)
by: Kou, Genghua, et al.
Published: (2025)
DriveAnchor: Progressive Anchor-based Flow Learning for Autonomous Driving Planning
by: Yan, Limin, et al.
Published: (2026)
by: Yan, Limin, et al.
Published: (2026)
STAR: A First-Ever Dataset and A Large-Scale Benchmark for Scene Graph Generation in Large-Size Satellite Imagery
by: Li, Yansheng, et al.
Published: (2024)
by: Li, Yansheng, et al.
Published: (2024)
GraphPilot: Grounded Scene Graph Conditioning for Language-Based Autonomous Driving
by: Schmidt, Fabian, et al.
Published: (2025)
by: Schmidt, Fabian, et al.
Published: (2025)
SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving
by: You, Zihan, et al.
Published: (2026)
by: You, Zihan, et al.
Published: (2026)
RaFD: Flow-Guided Radar Detection for Robust Autonomous Driving
by: Yang, Shuocheng, et al.
Published: (2025)
by: Yang, Shuocheng, et al.
Published: (2025)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
by: Wang, Shihao, et al.
Published: (2024)
by: Wang, Shihao, et al.
Published: (2024)
DriveCode: Domain Specific Numerical Encoding for LLM-Based Autonomous Driving
by: Wang, Zhiye, et al.
Published: (2026)
by: Wang, Zhiye, et al.
Published: (2026)
AutoTrust: Benchmarking Trustworthiness in Large Vision Language Models for Autonomous Driving
by: Xing, Shuo, et al.
Published: (2024)
by: Xing, Shuo, et al.
Published: (2024)
ROVR-Open-Dataset: A Large-Scale Depth Dataset for Autonomous Driving
by: Guo, Xianda, et al.
Published: (2025)
by: Guo, Xianda, et al.
Published: (2025)
doScenes: An Autonomous Driving Dataset with Natural Language Instruction for Human Interaction and Vision-Language Navigation
by: Roy, Parthib, et al.
Published: (2024)
by: Roy, Parthib, et al.
Published: (2024)
S2R-HDR: A Large-Scale Rendered Dataset for HDR Fusion
by: Wang, Yujin, et al.
Published: (2025)
by: Wang, Yujin, et al.
Published: (2025)
Driving with A Thousand Faces: A Benchmark for Closed-Loop Personalized End-to-End Autonomous Driving
by: Dong, Xiaoru, et al.
Published: (2026)
by: Dong, Xiaoru, et al.
Published: (2026)
Glad: A Streaming Scene Generator for Autonomous Driving
by: Xie, Bin, et al.
Published: (2025)
by: Xie, Bin, et al.
Published: (2025)
Formal Synthesis of Controllers for Safety-Critical Autonomous Systems: Developments and Challenges
by: Yin, Xiang, et al.
Published: (2024)
by: Yin, Xiang, et al.
Published: (2024)
Vision-Based Natural Language Scene Understanding for Autonomous Driving: An Extended Dataset and a New Model for Traffic Scene Description Generation
by: Zadeh, Danial Sadrian, et al.
Published: (2026)
by: Zadeh, Danial Sadrian, et al.
Published: (2026)
BEV-TSR: Text-Scene Retrieval in BEV Space for Autonomous Driving
by: Tang, Tao, et al.
Published: (2024)
by: Tang, Tao, et al.
Published: (2024)
An Instance-Centric Panoptic Occupancy Prediction Benchmark for Autonomous Driving
by: Feng, Yi, et al.
Published: (2026)
by: Feng, Yi, et al.
Published: (2026)
DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
by: Tian, Xiaoyu, et al.
Published: (2024)
by: Tian, Xiaoyu, et al.
Published: (2024)
A Generalized Control Revision Method for Autonomous Driving Safety
by: Zhu, Zehang, et al.
Published: (2024)
by: Zhu, Zehang, et al.
Published: (2024)
Similar Items
-
RAC3: Retrieval-Augmented Corner Case Comprehension for Autonomous Driving with Vision-Language Models
by: Wang, Yujin, et al.
Published: (2024) -
RAD: Retrieval-Augmented Decision-Making of Meta-Actions with Vision-Language Models in Autonomous Driving
by: Wang, Yujin, et al.
Published: (2025) -
ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driving
by: Ruan, Qiyu, et al.
Published: (2026) -
KEPT: Knowledge-Enhanced Prediction of Trajectories from Consecutive Driving Frames with Vision-Language Models
by: Wang, Yujin, et al.
Published: (2025) -
BIDA: A Bi-level Interaction Decision-making Algorithm for Autonomous Vehicles in Dynamic Traffic Scenarios
by: Yu, Liyang, et al.
Published: (2025)