RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yan, Tianyi, Han, Wencheng, Zhou, Xia, Zhang, Xueyang, Zhan, Kun, Xu, Cheng-zhong, Shen, Jianbing
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911228959391744
author Yan, Tianyi
Han, Wencheng
Zhou, Xia
Zhang, Xueyang
Zhan, Kun
Xu, Cheng-zhong
Shen, Jianbing
author_facet Yan, Tianyi
Han, Wencheng
Zhou, Xia
Zhang, Xueyang
Zhan, Kun
Xu, Cheng-zhong
Shen, Jianbing
contents Synthetic data is crucial for advancing autonomous driving (AD) systems, yet current state-of-the-art video generation models, despite their visual realism, suffer from subtle geometric distortions that limit their utility for downstream perception tasks. We identify and quantify this critical issue, demonstrating a significant performance gap in 3D object detection when using synthetic versus real data. To address this, we introduce Reinforcement Learning with Geometric Feedback (RLGF), RLGF uniquely refines video diffusion models by incorporating rewards from specialized latent-space AD perception models. Its core components include an efficient Latent-Space Windowing Optimization technique for targeted feedback during diffusion, and a Hierarchical Geometric Reward (HGR) system providing multi-level rewards for point-line-plane alignment, and scene occupancy coherence. To quantify these distortions, we propose GeoScores. Applied to models like DiVE on nuScenes, RLGF substantially reduces geometric errors (e.g., VP error by 21\%, Depth error by 57\%) and dramatically improves 3D object detection mAP by 12.7\%, narrowing the gap to real-data performance. RLGF offers a plug-and-play solution for generating geometrically sound and reliable synthetic videos for AD development.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16500
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation
Yan, Tianyi
Han, Wencheng
Zhou, Xia
Zhang, Xueyang
Zhan, Kun
Xu, Cheng-zhong
Shen, Jianbing
Computer Vision and Pattern Recognition
Synthetic data is crucial for advancing autonomous driving (AD) systems, yet current state-of-the-art video generation models, despite their visual realism, suffer from subtle geometric distortions that limit their utility for downstream perception tasks. We identify and quantify this critical issue, demonstrating a significant performance gap in 3D object detection when using synthetic versus real data. To address this, we introduce Reinforcement Learning with Geometric Feedback (RLGF), RLGF uniquely refines video diffusion models by incorporating rewards from specialized latent-space AD perception models. Its core components include an efficient Latent-Space Windowing Optimization technique for targeted feedback during diffusion, and a Hierarchical Geometric Reward (HGR) system providing multi-level rewards for point-line-plane alignment, and scene occupancy coherence. To quantify these distortions, we propose GeoScores. Applied to models like DiVE on nuScenes, RLGF substantially reduces geometric errors (e.g., VP error by 21\%, Depth error by 57\%) and dramatically improves 3D object detection mAP by 12.7\%, narrowing the gap to real-data performance. RLGF offers a plug-and-play solution for generating geometrically sound and reliable synthetic videos for AD development.
title RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.16500