Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Lijin, Huang, Jianing, Huang, Zhongzhan, Liu, Shu, Yang, Hao
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2604.27366
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914519941382144
author Yang, Lijin
Huang, Jianing
Huang, Zhongzhan
Liu, Shu
Yang, Hao
author_facet Yang, Lijin
Huang, Jianing
Huang, Zhongzhan
Liu, Shu
Yang, Hao
contents Recent advances in vision language action (VLA) models have shown remarkable potential for autonomous driving by directly mapping multimodal inputs to control signals. However, previous VLA-based methods have not explicitly exploited the critic capability of VLAs to refine driving decisions, even though such capability has been well demonstrated in other LLM-based domains, thereby limiting their performance in complex closed-loop scenarios. In this work, we present a theoretically inspired two-stage framework, CriticVLA, which extends the role of VLAs from acting to judging. CriticVLA first generates a rough trajectory and then refines it through multimodal evaluation and single-step optimization guided by a VLA-based critic, yielding higher-quality driving behaviors. To support this process, we construct a large-scale synthetic dataset of 12.9 million annotated trajectories covering diverse driving scenarios, which enhances the critic's reasoning and refinement abilities. Extensive closed-loop experiments on the Bench2Drive benchmark show that CriticVLA significantly surpasses state-of-the-art baselines, achieving a 73.33% total success rate and delivering about 30% improvement in challenging scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27366
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Judge, Then Drive: A Critic-Centric Vision Language Action Framework for Autonomous Driving
Yang, Lijin
Huang, Jianing
Huang, Zhongzhan
Liu, Shu
Yang, Hao
Computer Vision and Pattern Recognition
Recent advances in vision language action (VLA) models have shown remarkable potential for autonomous driving by directly mapping multimodal inputs to control signals. However, previous VLA-based methods have not explicitly exploited the critic capability of VLAs to refine driving decisions, even though such capability has been well demonstrated in other LLM-based domains, thereby limiting their performance in complex closed-loop scenarios. In this work, we present a theoretically inspired two-stage framework, CriticVLA, which extends the role of VLAs from acting to judging. CriticVLA first generates a rough trajectory and then refines it through multimodal evaluation and single-step optimization guided by a VLA-based critic, yielding higher-quality driving behaviors. To support this process, we construct a large-scale synthetic dataset of 12.9 million annotated trajectories covering diverse driving scenarios, which enhances the critic's reasoning and refinement abilities. Extensive closed-loop experiments on the Bench2Drive benchmark show that CriticVLA significantly surpasses state-of-the-art baselines, achieving a 73.33% total success rate and delivering about 30% improvement in challenging scenarios.
title Judge, Then Drive: A Critic-Centric Vision Language Action Framework for Autonomous Driving
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.27366