Generative Scenario Rollouts for End-to-End Autonomous Driving

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yasarla, Rajeev, Hegde, Deepti, Han, Shizhong, Cheng, Hsin-Pai, Shi, Yunxiao, Sadeghigooghari, Meysam, Mahajan, Shweta, Bhattacharyya, Apratim, Liu, Litian, Garrepalli, Risheek, Svantesson, Thomas, Porikli, Fatih, Cai, Hong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909992624324608
author Yasarla, Rajeev
Hegde, Deepti
Han, Shizhong
Cheng, Hsin-Pai
Shi, Yunxiao
Sadeghigooghari, Meysam
Mahajan, Shweta
Bhattacharyya, Apratim
Liu, Litian
Garrepalli, Risheek
Svantesson, Thomas
Porikli, Fatih
Cai, Hong
author_facet Yasarla, Rajeev
Hegde, Deepti
Han, Shizhong
Cheng, Hsin-Pai
Shi, Yunxiao
Sadeghigooghari, Meysam
Mahajan, Shweta
Bhattacharyya, Apratim
Liu, Litian
Garrepalli, Risheek
Svantesson, Thomas
Porikli, Fatih
Cai, Hong
contents Vision-Language-Action (VLA) models are emerging as highly effective planning models for end-to-end autonomous driving systems. However, current works mostly rely on imitation learning from sparse trajectory annotations and under-utilize their potential as generative models. We propose Generative Scenario Rollouts (GeRo), a plug-and-play framework for VLA models that jointly performs planning and generation of language-grounded future traffic scenes through an autoregressive rollout strategy. First, a VLA model is trained to encode ego vehicle and agent dynamics into latent tokens under supervision from planning, motion, and language tasks, facilitating text-aligned generation. Next, GeRo performs language-conditioned autoregressive generation. Given multi-view images, a scenario description, and ego-action questions, it generates future latent tokens and textual responses to guide long-horizon rollouts. A rollout-consistency loss stabilizes predictions using ground truth or pseudo-labels, mitigating drift and preserving text-action alignment. This design enables GeRo to perform temporally consistent, language-grounded rollouts that support long-horizon reasoning and multi-agent planning. On Bench2Drive, GeRo improves driving score and success rate by +15.7 and +26.2, respectively. By integrating reinforcement learning with generative rollouts, GeRo achieves state-of-the-art closed-loop and open-loop performance, demonstrating strong zero-shot robustness. These results highlight the promise of generative, language-conditioned reasoning as a foundation for safer and more interpretable end-to-end autonomous driving.
format Preprint
id arxiv_https___arxiv_org_abs_2601_11475
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Generative Scenario Rollouts for End-to-End Autonomous Driving
Yasarla, Rajeev
Hegde, Deepti
Han, Shizhong
Cheng, Hsin-Pai
Shi, Yunxiao
Sadeghigooghari, Meysam
Mahajan, Shweta
Bhattacharyya, Apratim
Liu, Litian
Garrepalli, Risheek
Svantesson, Thomas
Porikli, Fatih
Cai, Hong
Computer Vision and Pattern Recognition
Vision-Language-Action (VLA) models are emerging as highly effective planning models for end-to-end autonomous driving systems. However, current works mostly rely on imitation learning from sparse trajectory annotations and under-utilize their potential as generative models. We propose Generative Scenario Rollouts (GeRo), a plug-and-play framework for VLA models that jointly performs planning and generation of language-grounded future traffic scenes through an autoregressive rollout strategy. First, a VLA model is trained to encode ego vehicle and agent dynamics into latent tokens under supervision from planning, motion, and language tasks, facilitating text-aligned generation. Next, GeRo performs language-conditioned autoregressive generation. Given multi-view images, a scenario description, and ego-action questions, it generates future latent tokens and textual responses to guide long-horizon rollouts. A rollout-consistency loss stabilizes predictions using ground truth or pseudo-labels, mitigating drift and preserving text-action alignment. This design enables GeRo to perform temporally consistent, language-grounded rollouts that support long-horizon reasoning and multi-agent planning. On Bench2Drive, GeRo improves driving score and success rate by +15.7 and +26.2, respectively. By integrating reinforcement learning with generative rollouts, GeRo achieves state-of-the-art closed-loop and open-loop performance, demonstrating strong zero-shot robustness. These results highlight the promise of generative, language-conditioned reasoning as a foundation for safer and more interpretable end-to-end autonomous driving.
title Generative Scenario Rollouts for End-to-End Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.11475