UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866915666625298432 |
|---|---|
| author | Lu, Hao Liu, Ziyang Jiang, Guangfeng Luo, Yuanfei Chen, Sheng Zhang, Yangang Chen, Ying-Cong |
| author_facet | Lu, Hao Liu, Ziyang Jiang, Guangfeng Luo, Yuanfei Chen, Sheng Zhang, Yangang Chen, Ying-Cong |
| contents | Autonomous driving (AD) systems struggle in long-tail scenarios due to limited world knowledge and weak visual dynamic modeling. Existing vision-language-action (VLA)-based methods cannot leverage unlabeled videos for visual causal learning, while world model-based methods lack reasoning capabilities from large language models. In this paper, we construct multiple specialized datasets providing reasoning and planning annotations for complex scenarios. Then, a unified Understanding-Generation-Planning framework, named UniUGP, is proposed to synergize scene reasoning, future video generation, and trajectory planning through a hybrid expert architecture. By integrating pre-trained VLMs and video generation models, UniUGP leverages visual dynamics and semantic reasoning to enhance planning performance. Taking multi-frame observations and language instructions as input, it produces interpretable chain-of-thought reasoning, physically consistent trajectories, and coherent future videos. We introduce a four-stage training strategy that progressively builds these capabilities across multiple existing AD datasets, along with the proposed specialized datasets. Experiments demonstrate state-of-the-art performance in perception, reasoning, and decision-making, with superior generalization to challenging long-tail situations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_09864 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving Lu, Hao Liu, Ziyang Jiang, Guangfeng Luo, Yuanfei Chen, Sheng Zhang, Yangang Chen, Ying-Cong Computer Vision and Pattern Recognition Autonomous driving (AD) systems struggle in long-tail scenarios due to limited world knowledge and weak visual dynamic modeling. Existing vision-language-action (VLA)-based methods cannot leverage unlabeled videos for visual causal learning, while world model-based methods lack reasoning capabilities from large language models. In this paper, we construct multiple specialized datasets providing reasoning and planning annotations for complex scenarios. Then, a unified Understanding-Generation-Planning framework, named UniUGP, is proposed to synergize scene reasoning, future video generation, and trajectory planning through a hybrid expert architecture. By integrating pre-trained VLMs and video generation models, UniUGP leverages visual dynamics and semantic reasoning to enhance planning performance. Taking multi-frame observations and language instructions as input, it produces interpretable chain-of-thought reasoning, physically consistent trajectories, and coherent future videos. We introduce a four-stage training strategy that progressively builds these capabilities across multiple existing AD datasets, along with the proposed specialized datasets. Experiments demonstrate state-of-the-art performance in perception, reasoning, and decision-making, with superior generalization to challenging long-tail situations. |
| title | UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.09864 |