ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ma, Yunsheng, Yaman, Burhaneddin, Ye, Xin, Yurt, Mahmut, Luo, Jingru, Mallik, Abhirup, Wang, Ziran, Ren, Liu
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915295764938752
author Ma, Yunsheng
Yaman, Burhaneddin
Ye, Xin
Yurt, Mahmut
Luo, Jingru
Mallik, Abhirup
Wang, Ziran
Ren, Liu
author_facet Ma, Yunsheng
Yaman, Burhaneddin
Ye, Xin
Yurt, Mahmut
Luo, Jingru
Mallik, Abhirup
Wang, Ziran
Ren, Liu
contents Recent advances have explored integrating large language models (LLMs) into end-to-end autonomous driving systems to enhance generalization and interpretability. However, most existing approaches are limited to either driving performance or vision-language reasoning, making it difficult to achieve both simultaneously. In this paper, we propose ALN-P3, a unified co-distillation framework that introduces cross-modal alignment between "fast" vision-based autonomous driving systems and "slow" language-driven reasoning modules. ALN-P3 incorporates three novel alignment mechanisms: Perception Alignment (P1A), Prediction Alignment (P2A), and Planning Alignment (P3A), which explicitly align visual tokens with corresponding linguistic outputs across the full perception, prediction, and planning stack. All alignment modules are applied only during training and incur no additional costs during inference. Extensive experiments on four challenging benchmarks-nuScenes, Nu-X, TOD3Cap, and nuScenes QA-demonstrate that ALN-P3 significantly improves both driving decisions and language reasoning, achieving state-of-the-art results.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15158
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving
Ma, Yunsheng
Yaman, Burhaneddin
Ye, Xin
Yurt, Mahmut
Luo, Jingru
Mallik, Abhirup
Wang, Ziran
Ren, Liu
Computer Vision and Pattern Recognition
Computation and Language
Recent advances have explored integrating large language models (LLMs) into end-to-end autonomous driving systems to enhance generalization and interpretability. However, most existing approaches are limited to either driving performance or vision-language reasoning, making it difficult to achieve both simultaneously. In this paper, we propose ALN-P3, a unified co-distillation framework that introduces cross-modal alignment between "fast" vision-based autonomous driving systems and "slow" language-driven reasoning modules. ALN-P3 incorporates three novel alignment mechanisms: Perception Alignment (P1A), Prediction Alignment (P2A), and Planning Alignment (P3A), which explicitly align visual tokens with corresponding linguistic outputs across the full perception, prediction, and planning stack. All alignment modules are applied only during training and incur no additional costs during inference. Extensive experiments on four challenging benchmarks-nuScenes, Nu-X, TOD3Cap, and nuScenes QA-demonstrate that ALN-P3 significantly improves both driving decisions and language reasoning, achieving state-of-the-art results.
title ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2505.15158