TrajDiff: End-to-end Autonomous Driving without Perception Annotation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gui, Xingtai, Zhao, Jianbo, Han, Wencheng, Wang, Jikai, Gong, Jiahao, Tan, Feiyang, Xu, Cheng-zhong, Shen, Jianbing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917114336509952
author Gui, Xingtai
Zhao, Jianbo
Han, Wencheng
Wang, Jikai
Gong, Jiahao
Tan, Feiyang
Xu, Cheng-zhong
Shen, Jianbing
author_facet Gui, Xingtai
Zhao, Jianbo
Han, Wencheng
Wang, Jikai
Gong, Jiahao
Tan, Feiyang
Xu, Cheng-zhong
Shen, Jianbing
contents End-to-end autonomous driving systems directly generate driving policies from raw sensor inputs. While these systems can extract effective environmental features for planning, relying on auxiliary perception tasks, developing perception annotation-free planning paradigms has become increasingly critical due to the high cost of manual perception annotation. In this work, we propose TrajDiff, a Trajectory-oriented BEV Conditioned Diffusion framework that establishes a fully perception annotation-free generative method for end-to-end autonomous driving. TrajDiff requires only raw sensor inputs and future trajectory, constructing Gaussian BEV heatmap targets that inherently capture driving modalities. We design a simple yet effective trajectory-oriented BEV encoder to extract the TrajBEV feature without perceptual supervision. Furthermore, we introduce Trajectory-oriented BEV Diffusion Transformer (TB-DiT), which leverages ego-state information and the predicted TrajBEV features to directly generate diverse yet plausible trajectories, eliminating the need for handcrafted motion priors. Beyond architectural innovations, TrajDiff enables exploration of data scaling benefits in the annotation-free setting. Evaluated on the NAVSIM benchmark, TrajDiff achieves 87.5 PDMS, establishing state-of-the-art performance among all annotation-free methods. With data scaling, it further improves to 88.5 PDMS, which is comparable to advanced perception-based approaches. Our code and model will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2512_00723
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TrajDiff: End-to-end Autonomous Driving without Perception Annotation
Gui, Xingtai
Zhao, Jianbo
Han, Wencheng
Wang, Jikai
Gong, Jiahao
Tan, Feiyang
Xu, Cheng-zhong
Shen, Jianbing
Computer Vision and Pattern Recognition
Robotics
End-to-end autonomous driving systems directly generate driving policies from raw sensor inputs. While these systems can extract effective environmental features for planning, relying on auxiliary perception tasks, developing perception annotation-free planning paradigms has become increasingly critical due to the high cost of manual perception annotation. In this work, we propose TrajDiff, a Trajectory-oriented BEV Conditioned Diffusion framework that establishes a fully perception annotation-free generative method for end-to-end autonomous driving. TrajDiff requires only raw sensor inputs and future trajectory, constructing Gaussian BEV heatmap targets that inherently capture driving modalities. We design a simple yet effective trajectory-oriented BEV encoder to extract the TrajBEV feature without perceptual supervision. Furthermore, we introduce Trajectory-oriented BEV Diffusion Transformer (TB-DiT), which leverages ego-state information and the predicted TrajBEV features to directly generate diverse yet plausible trajectories, eliminating the need for handcrafted motion priors. Beyond architectural innovations, TrajDiff enables exploration of data scaling benefits in the annotation-free setting. Evaluated on the NAVSIM benchmark, TrajDiff achieves 87.5 PDMS, establishing state-of-the-art performance among all annotation-free methods. With data scaling, it further improves to 88.5 PDMS, which is comparable to advanced perception-based approaches. Our code and model will be made publicly available.
title TrajDiff: End-to-end Autonomous Driving without Perception Annotation
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2512.00723