Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Zhenghao "Mark", Ding, Wenhao, You, Yurong, Chen, Yuxiao, Luo, Wenjie, Tian, Thomas, Cao, Yulong, Sharma, Apoorva, Xu, Danfei, Ivanovic, Boris, Li, Boyi, Zhou, Bolei, Wang, Yan, Pavone, Marco
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918266269597696
author Peng, Zhenghao "Mark"
Ding, Wenhao
You, Yurong
Chen, Yuxiao
Luo, Wenjie
Tian, Thomas
Cao, Yulong
Sharma, Apoorva
Xu, Danfei
Ivanovic, Boris
Li, Boyi
Zhou, Bolei
Wang, Yan
Pavone, Marco
author_facet Peng, Zhenghao "Mark"
Ding, Wenhao
You, Yurong
Chen, Yuxiao
Luo, Wenjie
Tian, Thomas
Cao, Yulong
Sharma, Apoorva
Xu, Danfei
Ivanovic, Boris
Li, Boyi
Zhou, Bolei
Wang, Yan
Pavone, Marco
contents Recent reasoning-augmented Vision-Language-Action (VLA) models have improved the interpretability of end-to-end autonomous driving by generating intermediate reasoning traces. Yet these models primarily describe what they perceive and intend to do, rarely questioning whether their planned actions are safe or appropriate. This work introduces Counterfactual VLA (CF-VLA), a self-reflective VLA framework that enables the model to reason about and revise its planned actions before execution. CF-VLA first generates time-segmented meta-actions that summarize driving intent, and then performs counterfactual reasoning conditioned on both the meta-actions and the visual context. This step simulates potential outcomes, identifies unsafe behaviors, and outputs corrected meta-actions that guide the final trajectory generation. To efficiently obtain such self-reflective capabilities, we propose a rollout-filter-label pipeline that mines high-value scenes from a base (non-counterfactual) VLA's rollouts and labels counterfactual reasoning traces for subsequent training rounds. Experiments on large-scale driving datasets show that CF-VLA improves trajectory accuracy by up to 17.6%, enhances safety metrics by 20.5%, and exhibits adaptive thinking: it only enables counterfactual reasoning in challenging scenarios. By transforming reasoning traces from one-shot descriptions to causal self-correction signals, CF-VLA takes a step toward self-reflective autonomous driving agents that learn to think before they act.
format Preprint
id arxiv_https___arxiv_org_abs_2512_24426
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
Peng, Zhenghao "Mark"
Ding, Wenhao
You, Yurong
Chen, Yuxiao
Luo, Wenjie
Tian, Thomas
Cao, Yulong
Sharma, Apoorva
Xu, Danfei
Ivanovic, Boris
Li, Boyi
Zhou, Bolei
Wang, Yan
Pavone, Marco
Robotics
Recent reasoning-augmented Vision-Language-Action (VLA) models have improved the interpretability of end-to-end autonomous driving by generating intermediate reasoning traces. Yet these models primarily describe what they perceive and intend to do, rarely questioning whether their planned actions are safe or appropriate. This work introduces Counterfactual VLA (CF-VLA), a self-reflective VLA framework that enables the model to reason about and revise its planned actions before execution. CF-VLA first generates time-segmented meta-actions that summarize driving intent, and then performs counterfactual reasoning conditioned on both the meta-actions and the visual context. This step simulates potential outcomes, identifies unsafe behaviors, and outputs corrected meta-actions that guide the final trajectory generation. To efficiently obtain such self-reflective capabilities, we propose a rollout-filter-label pipeline that mines high-value scenes from a base (non-counterfactual) VLA's rollouts and labels counterfactual reasoning traces for subsequent training rounds. Experiments on large-scale driving datasets show that CF-VLA improves trajectory accuracy by up to 17.6%, enhances safety metrics by 20.5%, and exhibits adaptive thinking: it only enables counterfactual reasoning in challenging scenarios. By transforming reasoning traces from one-shot descriptions to causal self-correction signals, CF-VLA takes a step toward self-reflective autonomous driving agents that learn to think before they act.
title Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
topic Robotics
url https://arxiv.org/abs/2512.24426