Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Oksuz, Kemal, Buburuzan, Alexandru, Knittel, Anthony, Yao, Yuhan, Dokania, Puneet K.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917357480312832
author Oksuz, Kemal
Buburuzan, Alexandru
Knittel, Anthony
Yao, Yuhan
Dokania, Puneet K.
author_facet Oksuz, Kemal
Buburuzan, Alexandru
Knittel, Anthony
Yao, Yuhan
Dokania, Puneet K.
contents The emergence of multi-modal foundation models has markedly transformed the technology for autonomous driving, shifting away from conventional and mostly hand-crafted design choices towards unified, foundation-model-based approaches, capable of directly inferring motion trajectories from raw sensory inputs. This new class of methods can also incorporate natural language as an additional modality, with Vision-Language-Action (VLA) models serving as a representative example. In this review, we provide a comprehensive examination of such methods through a unifying taxonomy to critically evaluate their architectural design choices, methodological strengths, and their inherent capabilities and limitations. Our survey covers 37 recently proposed approaches that span the landscape of trajectory planning with foundation models. Furthermore, we assess these approaches with respect to the openness of their source code and datasets, offering valuable information to practitioners and researchers. We provide an accompanying webpage that catalogues the methods based on our taxonomy, available at: https://github.com/fiveai/FMs-for-driving-trajectories
format Preprint
id arxiv_https___arxiv_org_abs_2512_00021
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
Oksuz, Kemal
Buburuzan, Alexandru
Knittel, Anthony
Yao, Yuhan
Dokania, Puneet K.
Robotics
Computer Vision and Pattern Recognition
The emergence of multi-modal foundation models has markedly transformed the technology for autonomous driving, shifting away from conventional and mostly hand-crafted design choices towards unified, foundation-model-based approaches, capable of directly inferring motion trajectories from raw sensory inputs. This new class of methods can also incorporate natural language as an additional modality, with Vision-Language-Action (VLA) models serving as a representative example. In this review, we provide a comprehensive examination of such methods through a unifying taxonomy to critically evaluate their architectural design choices, methodological strengths, and their inherent capabilities and limitations. Our survey covers 37 recently proposed approaches that span the landscape of trajectory planning with foundation models. Furthermore, we assess these approaches with respect to the openness of their source code and datasets, offering valuable information to practitioners and researchers. We provide an accompanying webpage that catalogues the methods based on our taxonomy, available at: https://github.com/fiveai/FMs-for-driving-trajectories
title Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.00021