WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Pengxuan, Lu, Ben, Xia, Zhongpu, Han, Chao, Gao, Yinfeng, Zhang, Teng, Zhan, Kun, Lang, XianPeng, Zheng, Yupeng, Zhang, Qichao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911332036509696
author Yang, Pengxuan
Lu, Ben
Xia, Zhongpu
Han, Chao
Gao, Yinfeng
Zhang, Teng
Zhan, Kun
Lang, XianPeng
Zheng, Yupeng
Zhang, Qichao
author_facet Yang, Pengxuan
Lu, Ben
Xia, Zhongpu
Han, Chao
Gao, Yinfeng
Zhang, Teng
Zhan, Kun
Lang, XianPeng
Zheng, Yupeng
Zhang, Qichao
contents Latent World Models enhance scene representation through temporal self-supervised learning, presenting a perception annotation-free paradigm for end-to-end autonomous driving. However, the reconstruction-oriented representation learning tangles perception with planning tasks, leading to suboptimal optimization for planning. To address this challenge, we propose WorldRFT, a planning-oriented latent world model framework that aligns scene representation learning with planning via a hierarchical planning decomposition and local-aware interactive refinement mechanism, augmented by reinforcement learning fine-tuning (RFT) to enhance safety-critical policy performance. Specifically, WorldRFT integrates a vision-geometry foundation model to improve 3D spatial awareness, employs hierarchical planning task decomposition to guide representation optimization, and utilizes local-aware iterative refinement to derive a planning-oriented driving policy. Furthermore, we introduce Group Relative Policy Optimization (GRPO), which applies trajectory Gaussianization and collision-aware rewards to fine-tune the driving policy, yielding systematic improvements in safety. WorldRFT achieves state-of-the-art (SOTA) performance on both open-loop nuScenes and closed-loop NavSim benchmarks. On nuScenes, it reduces collision rates by 83% (0.30% -> 0.05%). On NavSim, using camera-only sensors input, it attains competitive performance with the LiDAR-based SOTA method DiffusionDrive (87.8 vs. 88.1 PDMS).
format Preprint
id arxiv_https___arxiv_org_abs_2512_19133
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving
Yang, Pengxuan
Lu, Ben
Xia, Zhongpu
Han, Chao
Gao, Yinfeng
Zhang, Teng
Zhan, Kun
Lang, XianPeng
Zheng, Yupeng
Zhang, Qichao
Robotics
Computer Vision and Pattern Recognition
Latent World Models enhance scene representation through temporal self-supervised learning, presenting a perception annotation-free paradigm for end-to-end autonomous driving. However, the reconstruction-oriented representation learning tangles perception with planning tasks, leading to suboptimal optimization for planning. To address this challenge, we propose WorldRFT, a planning-oriented latent world model framework that aligns scene representation learning with planning via a hierarchical planning decomposition and local-aware interactive refinement mechanism, augmented by reinforcement learning fine-tuning (RFT) to enhance safety-critical policy performance. Specifically, WorldRFT integrates a vision-geometry foundation model to improve 3D spatial awareness, employs hierarchical planning task decomposition to guide representation optimization, and utilizes local-aware iterative refinement to derive a planning-oriented driving policy. Furthermore, we introduce Group Relative Policy Optimization (GRPO), which applies trajectory Gaussianization and collision-aware rewards to fine-tune the driving policy, yielding systematic improvements in safety. WorldRFT achieves state-of-the-art (SOTA) performance on both open-loop nuScenes and closed-loop NavSim benchmarks. On nuScenes, it reduces collision rates by 83% (0.30% -> 0.05%). On NavSim, using camera-only sensors input, it attains competitive performance with the LiDAR-based SOTA method DiffusionDrive (87.8 vs. 88.1 PDMS).
title WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.19133