Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Zhenghao "Mark", Ding, Wenhao, You, Yurong, Chen, Yuxiao, Luo, Wenjie, Tian, Thomas, Cao, Yulong, Sharma, Apoorva, Xu, Danfei, Ivanovic, Boris, Li, Boyi, Zhou, Bolei, Wang, Yan, Pavone, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes
by: Yang, Jiawei, et al.
Published: (2024)
by: Yang, Jiawei, et al.
Published: (2024)
DreamDrive: Generative 4D Scene Modeling from Street View Images
by: Mao, Jiageng, et al.
Published: (2024)
by: Mao, Jiageng, et al.
Published: (2024)
Accelerating Structured Chain-of-Thought in Autonomous Vehicles
by: Gu, Yi, et al.
Published: (2026)
by: Gu, Yi, et al.
Published: (2026)
The Case for Negative Data: From Crash Reports to Counterfactuals for Reasonable Driving
by: Patrikar, Jay, et al.
Published: (2025)
by: Patrikar, Jay, et al.
Published: (2025)
Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving
by: Ivanovic, Boris, et al.
Published: (2025)
by: Ivanovic, Boris, et al.
Published: (2025)
Latent Chain-of-Thought World Modeling for End-to-End Driving
by: Tan, Shuhan, et al.
Published: (2025)
by: Tan, Shuhan, et al.
Published: (2025)
LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation Models
by: Lu, Ziqi, et al.
Published: (2024)
by: Lu, Ziqi, et al.
Published: (2024)
Promptable Closed-loop Traffic Simulation
by: Tan, Shuhan, et al.
Published: (2024)
by: Tan, Shuhan, et al.
Published: (2024)
dVLM-AD: Enhance Diffusion Vision-Language-Model for Driving via Controllable Reasoning
by: Ma, Yingzi, et al.
Published: (2025)
by: Ma, Yingzi, et al.
Published: (2025)
Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
by: Yang, Jiawei, et al.
Published: (2025)
by: Yang, Jiawei, et al.
Published: (2025)
Closed-Loop Supervised Fine-Tuning of Tokenized Traffic Models
by: Zhang, Zhejun, et al.
Published: (2024)
by: Zhang, Zhejun, et al.
Published: (2024)
Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving
by: Tian, Ran, et al.
Published: (2024)
by: Tian, Ran, et al.
Published: (2024)
Safety Evaluation of Motion Plans Using Trajectory Predictors as Forward Reachable Set Estimators
by: Chakraborty, Kaustav, et al.
Published: (2025)
by: Chakraborty, Kaustav, et al.
Published: (2025)
Language-Image Models with 3D Understanding
by: Cho, Jang Hyun, et al.
Published: (2024)
by: Cho, Jang Hyun, et al.
Published: (2024)
System-Level Safety Monitoring and Recovery for Perception Failures in Autonomous Vehicles
by: Chakraborty, Kaustav, et al.
Published: (2024)
by: Chakraborty, Kaustav, et al.
Published: (2024)
AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning
by: Zhou, Zewei, et al.
Published: (2025)
by: Zhou, Zewei, et al.
Published: (2025)
Gen-Drive: Enhancing Diffusion Generative Driving Policies with Reward Modeling and Reinforcement Learning Fine-tuning
by: Huang, Zhiyu, et al.
Published: (2024)
by: Huang, Zhiyu, et al.
Published: (2024)
RealDrive: Retrieval-Augmented Driving with Diffusion Models
by: Ding, Wenhao, et al.
Published: (2025)
by: Ding, Wenhao, et al.
Published: (2025)
OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning
by: Lin, Fanqi, et al.
Published: (2025)
by: Lin, Fanqi, et al.
Published: (2025)
Extrapolated Urban View Synthesis Benchmark
by: Han, Xiangyu, et al.
Published: (2024)
by: Han, Xiangyu, et al.
Published: (2024)
DTPP: Differentiable Joint Conditional Prediction and Cost Evaluation for Tree Policy Planning in Autonomous Driving
by: Huang, Zhiyu, et al.
Published: (2023)
by: Huang, Zhiyu, et al.
Published: (2023)
Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model
by: Li, Wenhao, et al.
Published: (2026)
by: Li, Wenhao, et al.
Published: (2026)
Driving Everywhere with Large Language Model Policy Adaptation
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
Robot-Gated Interactive Imitation Learning with Adaptive Intervention Mechanism
by: Cai, Haoyuan, et al.
Published: (2025)
by: Cai, Haoyuan, et al.
Published: (2025)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
by: Ding, Pengxiang, et al.
Published: (2023)
by: Ding, Pengxiang, et al.
Published: (2023)
RealGen: Retrieval Augmented Generation for Controllable Traffic Scenarios
by: Ding, Wenhao, et al.
Published: (2023)
by: Ding, Wenhao, et al.
Published: (2023)
FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos
by: Gan, Yulu, et al.
Published: (2025)
by: Gan, Yulu, et al.
Published: (2025)
StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement
by: Seo, Junwon, et al.
Published: (2026)
by: Seo, Junwon, et al.
Published: (2026)
Surprise Potential as a Measure of Interactivity in Driving Scenarios
by: Ding, Wenhao, et al.
Published: (2025)
by: Ding, Wenhao, et al.
Published: (2025)
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
by: Deng, Shengliang, et al.
Published: (2025)
by: Deng, Shengliang, et al.
Published: (2025)
VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
by: Wang, Chaoyang, et al.
Published: (2026)
by: Wang, Chaoyang, et al.
Published: (2026)
Producing and Leveraging Online Map Uncertainty in Trajectory Prediction
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
Accelerating Online Mapping and Behavior Prediction via Direct BEV Feature Attention
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
Learning from Active Human Involvement through Proxy Value Propagation
by: Peng, Zhenghao, et al.
Published: (2025)
by: Peng, Zhenghao, et al.
Published: (2025)
RoaD: Rollouts as Demonstrations for Closed-Loop Supervised Fine-Tuning of Autonomous Driving Policies
by: Garcia-Cobo, Guillermo, et al.
Published: (2025)
by: Garcia-Cobo, Guillermo, et al.
Published: (2025)
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
by: Sun, Jianli, et al.
Published: (2026)
by: Sun, Jianli, et al.
Published: (2026)
VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
by: Ye, Angen, et al.
Published: (2025)
by: Ye, Angen, et al.
Published: (2025)
CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human
by: Sun, Nan, et al.
Published: (2025)
by: Sun, Nan, et al.
Published: (2025)
InstantSplat: Sparse-view Gaussian Splatting in Seconds
by: Fan, Zhiwen, et al.
Published: (2024)
by: Fan, Zhiwen, et al.
Published: (2024)
Similar Items
-
STORM: Spatio-Temporal Reconstruction Model for Large-Scale Outdoor Scenes
by: Yang, Jiawei, et al.
Published: (2024) -
DreamDrive: Generative 4D Scene Modeling from Street View Images
by: Mao, Jiageng, et al.
Published: (2024) -
Accelerating Structured Chain-of-Thought in Autonomous Vehicles
by: Gu, Yi, et al.
Published: (2026) -
The Case for Negative Data: From Crash Reports to Counterfactuals for Reasonable Driving
by: Patrikar, Jay, et al.
Published: (2025) -
Efficient Multi-Camera Tokenization with Triplanes for End-to-End Driving
by: Ivanovic, Boris, et al.
Published: (2025)