ScenePilot-4K: A Large-Scale First-Person Dataset and Benchmark for Vision-Language Models in Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yujin, Zheng, Yutong, Fan, Wenxian, Wang, Tianyi, Chu, Hongqing, Zhang, Li, Gao, Bingzhao, Tian, Daxin, Wang, Jianqiang, Chen, Hong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908919950999552
author Wang, Yujin
Zheng, Yutong
Fan, Wenxian
Wang, Tianyi
Chu, Hongqing
Zhang, Li
Gao, Bingzhao
Tian, Daxin
Wang, Jianqiang
Chen, Hong
author_facet Wang, Yujin
Zheng, Yutong
Fan, Wenxian
Wang, Tianyi
Chu, Hongqing
Zhang, Li
Gao, Bingzhao
Tian, Daxin
Wang, Jianqiang
Chen, Hong
contents In this paper, we introduce ScenePilot-4K, a large-scale first-person dataset for safety-aware vision-language learning and evaluation in autonomous driving. Built from public online driving videos, ScenePilot-4K contains 3,847 hours of video and 27.7M front-view frames spanning 63 countries/regions and 1,210 cities. It jointly provides scene-level natural-language descriptions, risk assessment labels, key-participant annotations, ego trajectories, and camera parameters through a unified multi-stage annotation pipeline. Building on this dataset, we establish ScenePilot-Bench, a standardized benchmark that evaluates vision-language models along four complementary axes: scene understanding, spatial perception, motion planning, and GPT-based semantic alignment. The benchmark includes fine-grained metrics and geographic generalization settings that expose model robustness under cross-region and cross-traffic domain shifts. Baseline results on representative open-source and proprietary vision-language models show that current models remain competitive in high-level scene semantics but still exhibit substantial limitations in geometry-aware perception and planning-oriented reasoning. Beyond the released dataset itself, the proposed annotation pipeline serves as a reusable and extensible recipe for scalable dataset construction from public Internet driving videos. The codes and supplementary materials are available at: https://github.com/yjwangtj/ScenePilot-4K, with the dataset available at https://huggingface.co/datasets/larswangtj/ScenePilot-4K.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19582
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ScenePilot-4K: A Large-Scale First-Person Dataset and Benchmark for Vision-Language Models in Autonomous Driving
Wang, Yujin
Zheng, Yutong
Fan, Wenxian
Wang, Tianyi
Chu, Hongqing
Zhang, Li
Gao, Bingzhao
Tian, Daxin
Wang, Jianqiang
Chen, Hong
Computer Vision and Pattern Recognition
In this paper, we introduce ScenePilot-4K, a large-scale first-person dataset for safety-aware vision-language learning and evaluation in autonomous driving. Built from public online driving videos, ScenePilot-4K contains 3,847 hours of video and 27.7M front-view frames spanning 63 countries/regions and 1,210 cities. It jointly provides scene-level natural-language descriptions, risk assessment labels, key-participant annotations, ego trajectories, and camera parameters through a unified multi-stage annotation pipeline. Building on this dataset, we establish ScenePilot-Bench, a standardized benchmark that evaluates vision-language models along four complementary axes: scene understanding, spatial perception, motion planning, and GPT-based semantic alignment. The benchmark includes fine-grained metrics and geographic generalization settings that expose model robustness under cross-region and cross-traffic domain shifts. Baseline results on representative open-source and proprietary vision-language models show that current models remain competitive in high-level scene semantics but still exhibit substantial limitations in geometry-aware perception and planning-oriented reasoning. Beyond the released dataset itself, the proposed annotation pipeline serves as a reusable and extensible recipe for scalable dataset construction from public Internet driving videos. The codes and supplementary materials are available at: https://github.com/yjwangtj/ScenePilot-4K, with the dataset available at https://huggingface.co/datasets/larswangtj/ScenePilot-4K.
title ScenePilot-4K: A Large-Scale First-Person Dataset and Benchmark for Vision-Language Models in Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.19582