From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909983060262912 |
|---|---|
| author | Zheng, Huan Zhou, Yucheng Yan, Tianyi Su, Jiayi Chen, Hongjun Chen, Dubing Gui, Xingtai Han, Wencheng Tao, Runzhou Qiu, Zhongying Yang, Jianfei Shen, Jianbing |
| author_facet | Zheng, Huan Zhou, Yucheng Yan, Tianyi Su, Jiayi Chen, Hongjun Chen, Dubing Gui, Xingtai Han, Wencheng Tao, Runzhou Qiu, Zhongying Yang, Jianfei Shen, Jianbing |
| contents | While end-to-end autonomous driving has achieved remarkable progress in geometric control, current systems remain constrained by a command-following paradigm that relies on simple navigational instructions. Transitioning to genuinely intelligent agents requires the capability to interpret and fulfill high-level, abstract human intentions. However, this advancement is hindered by the lack of dedicated benchmarks and semantic-aware evaluation metrics. In this paper, we formally define the task of Intention-Driven End-to-End Autonomous Driving and present Intention-Drive, a comprehensive benchmark designed to bridge this gap. We construct a large-scale dataset featuring complex natural language intentions paired with high-fidelity sensor data. To overcome the limitations of conventional trajectory-based metrics, we introduce the Imagined Future Alignment (IFA), a novel evaluation protocol leveraging generative world models to assess the semantic fulfillment of human goals beyond mere geometric accuracy. Furthermore, we explore the solution space by proposing two distinct paradigms: an end-to-end vision-language planner and a hierarchical agent-based framework. The experiments reveal a critical dichotomy where existing models exhibit satisfactory driving stability but struggle significantly with intention fulfillment. Notably, the proposed frameworks demonstrate superior alignment with human intentions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_12302 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving Zheng, Huan Zhou, Yucheng Yan, Tianyi Su, Jiayi Chen, Hongjun Chen, Dubing Gui, Xingtai Han, Wencheng Tao, Runzhou Qiu, Zhongying Yang, Jianfei Shen, Jianbing Computer Vision and Pattern Recognition Computation and Language Robotics While end-to-end autonomous driving has achieved remarkable progress in geometric control, current systems remain constrained by a command-following paradigm that relies on simple navigational instructions. Transitioning to genuinely intelligent agents requires the capability to interpret and fulfill high-level, abstract human intentions. However, this advancement is hindered by the lack of dedicated benchmarks and semantic-aware evaluation metrics. In this paper, we formally define the task of Intention-Driven End-to-End Autonomous Driving and present Intention-Drive, a comprehensive benchmark designed to bridge this gap. We construct a large-scale dataset featuring complex natural language intentions paired with high-fidelity sensor data. To overcome the limitations of conventional trajectory-based metrics, we introduce the Imagined Future Alignment (IFA), a novel evaluation protocol leveraging generative world models to assess the semantic fulfillment of human goals beyond mere geometric accuracy. Furthermore, we explore the solution space by proposing two distinct paradigms: an end-to-end vision-language planner and a hierarchical agent-based framework. The experiments reveal a critical dichotomy where existing models exhibit satisfactory driving stability but struggle significantly with intention fulfillment. Notably, the proposed frameworks demonstrate superior alignment with human intentions. |
| title | From Human Intention to Action Prediction: Intention-Driven End-to-End Autonomous Driving |
| topic | Computer Vision and Pattern Recognition Computation and Language Robotics |
| url | https://arxiv.org/abs/2512.12302 |