When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fang, Yu, Feng, Yuchun, Jing, Dong, Liu, Jiaqi, Yang, Yue, Wei, Zhenyu, Szafir, Daniel, Ding, Mingyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BOSS: Benchmark for Observation Space Shift in Long-Horizon Task
von: Yang, Yue, et al.
Veröffentlicht: (2025)
von: Yang, Yue, et al.
Veröffentlicht: (2025)
ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis
von: Fang, Yu, et al.
Veröffentlicht: (2025)
von: Fang, Yu, et al.
Veröffentlicht: (2025)
LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies
von: Yang, Yue, et al.
Veröffentlicht: (2026)
von: Yang, Yue, et al.
Veröffentlicht: (2026)
Differentiate-and-Inject: Enhancing VLAs via Functional Differentiation Induced by In-Parameter Structural Reasoning
von: Hou, Jingyi, et al.
Veröffentlicht: (2026)
von: Hou, Jingyi, et al.
Veröffentlicht: (2026)
ARCADE: Scalable Demonstration Collection and Generation via Augmented Reality for Imitation Learning
von: Yang, Yue, et al.
Veröffentlicht: (2024)
von: Yang, Yue, et al.
Veröffentlicht: (2024)
One Hand to Rule Them All: Canonical Representations for Unified Dexterous Manipulation
von: Wei, Zhenyu, et al.
Veröffentlicht: (2026)
von: Wei, Zhenyu, et al.
Veröffentlicht: (2026)
Running VLAs at Real-time Speed
von: Ma, Yunchao, et al.
Veröffentlicht: (2025)
von: Ma, Yunchao, et al.
Veröffentlicht: (2025)
Realtime-VLA V2: Learning to Run VLAs Fast, Smooth, and Accurate
von: Yang, Chen, et al.
Veröffentlicht: (2026)
von: Yang, Chen, et al.
Veröffentlicht: (2026)
Augmented Reality Demonstrations for Scalable Robot Imitation Learning
von: Yang, Yue, et al.
Veröffentlicht: (2024)
von: Yang, Yue, et al.
Veröffentlicht: (2024)
FASTER: Rethinking Real-Time Flow VLAs
von: Lu, Yuxiang, et al.
Veröffentlicht: (2026)
von: Lu, Yuxiang, et al.
Veröffentlicht: (2026)
Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
von: Hancock, Asher J., et al.
Veröffentlicht: (2025)
von: Hancock, Asher J., et al.
Veröffentlicht: (2025)
Notes-to-Self: Scratchpad Augmented VLAs for Memory Dependent Manipulation Tasks
von: Haresh, Sanjay, et al.
Veröffentlicht: (2026)
von: Haresh, Sanjay, et al.
Veröffentlicht: (2026)
SITCOM: Scaling Inference-Time COMpute for VLAs
von: Saxena, Ayudh, et al.
Veröffentlicht: (2025)
von: Saxena, Ayudh, et al.
Veröffentlicht: (2025)
Shallow-π: Knowledge Distillation for Flow-based VLAs
von: Jeon, Boseong, et al.
Veröffentlicht: (2026)
von: Jeon, Boseong, et al.
Veröffentlicht: (2026)
Primitive Subspaces Mediate Few-Shot Transfer in VLAs
von: Singh, Anya, et al.
Veröffentlicht: (2026)
von: Singh, Anya, et al.
Veröffentlicht: (2026)
VLAs are Confined yet Capable of Generalizing to Novel Instructions
von: Li, Quanyi
Veröffentlicht: (2025)
von: Li, Quanyi
Veröffentlicht: (2025)
How Do VLAs Effectively Inherit from VLMs?
von: Zhang, Chuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Chuheng, et al.
Veröffentlicht: (2025)
How VLAs (Really) Work In Open-World Environments
von: Rasouli, Amir, et al.
Veröffentlicht: (2026)
von: Rasouli, Amir, et al.
Veröffentlicht: (2026)
Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs
von: Zhao, Jianchao, et al.
Veröffentlicht: (2026)
von: Zhao, Jianchao, et al.
Veröffentlicht: (2026)
Do World Action Models Generalize Better than VLAs? A Robustness Study
von: Zhang, Zhanguang, et al.
Veröffentlicht: (2026)
von: Zhang, Zhanguang, et al.
Veröffentlicht: (2026)
VLA-0: Building State-of-the-Art VLAs with Zero Modification
von: Goyal, Ankit, et al.
Veröffentlicht: (2025)
von: Goyal, Ankit, et al.
Veröffentlicht: (2025)
Deep Learning based Quasi-consciousness Training for Robot Intelligent Model
von: Li, Yuchun, et al.
Veröffentlicht: (2025)
von: Li, Yuchun, et al.
Veröffentlicht: (2025)
Language-Driven Policy Distillation for Cooperative Driving in Multi-Agent Reinforcement Learning
von: Liu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Liu, Jiaqi, et al.
Veröffentlicht: (2024)
Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
von: Peng, Zhenghao "Mark", et al.
Veröffentlicht: (2025)
von: Peng, Zhenghao "Mark", et al.
Veröffentlicht: (2025)
LaViRA: Language-Vision-Robot Actions Translation for Zero-Shot Vision Language Navigation in Continuous Environments
von: Ding, Hongyu, et al.
Veröffentlicht: (2025)
von: Ding, Hongyu, et al.
Veröffentlicht: (2025)
cVLA: Towards Efficient Camera-Space VLAs
von: Argus, Max, et al.
Veröffentlicht: (2025)
von: Argus, Max, et al.
Veröffentlicht: (2025)
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
von: Lin, Zijun, et al.
Veröffentlicht: (2025)
von: Lin, Zijun, et al.
Veröffentlicht: (2025)
Towards Interactive and Learnable Cooperative Driving Automation: a Large Language Model-Driven Decision-Making Framework
von: Fang, Shiyu, et al.
Veröffentlicht: (2024)
von: Fang, Shiyu, et al.
Veröffentlicht: (2024)
ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action Models
von: Fang, Zhou, et al.
Veröffentlicht: (2026)
von: Fang, Zhou, et al.
Veröffentlicht: (2026)
Integrated Hardware and Software Architecture for Industrial AGV with Manual Override Capability
von: Iob, Pietro, et al.
Veröffentlicht: (2024)
von: Iob, Pietro, et al.
Veröffentlicht: (2024)
How VLAs Fail Differently: Black-Box Action Monitoring Reveals Architecture-Specific Failure Signatures
von: Gupta, Krishnam
Veröffentlicht: (2026)
von: Gupta, Krishnam
Veröffentlicht: (2026)
Mixture of Horizons in Action Chunking
von: Jing, Dong, et al.
Veröffentlicht: (2025)
von: Jing, Dong, et al.
Veröffentlicht: (2025)
CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
von: Glossop, Catherine, et al.
Veröffentlicht: (2025)
von: Glossop, Catherine, et al.
Veröffentlicht: (2025)
CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation
von: Song, Kun, et al.
Veröffentlicht: (2025)
von: Song, Kun, et al.
Veröffentlicht: (2025)
Lost in Fog: Sensor Perturbations Expose Reasoning Fragility in Driving VLAs
von: Priyadershi, Abhinaw, et al.
Veröffentlicht: (2026)
von: Priyadershi, Abhinaw, et al.
Veröffentlicht: (2026)
Paper index: Designing an introductory HRI course (workshop at HRI 2024)
von: Admoni, Henny, et al.
Veröffentlicht: (2024)
von: Admoni, Henny, et al.
Veröffentlicht: (2024)
Scaling Sim-to-Real Reinforcement Learning for Robot VLAs with Generative 3D Worlds
von: Choi, Andrew, et al.
Veröffentlicht: (2026)
von: Choi, Andrew, et al.
Veröffentlicht: (2026)
Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation
von: Wei, Meng, et al.
Veröffentlicht: (2025)
von: Wei, Meng, et al.
Veröffentlicht: (2025)
What Frozen VLAs Already Know About Success: A Probing Study of Value-Like Structure in Foundation Robot Policies
von: Zhang, Jiachen, et al.
Veröffentlicht: (2026)
von: Zhang, Jiachen, et al.
Veröffentlicht: (2026)
VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
von: Yuan, Haoran, et al.
Veröffentlicht: (2026)
von: Yuan, Haoran, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BOSS: Benchmark for Observation Space Shift in Long-Horizon Task
von: Yang, Yue, et al.
Veröffentlicht: (2025) -
ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis
von: Fang, Yu, et al.
Veröffentlicht: (2025) -
LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies
von: Yang, Yue, et al.
Veröffentlicht: (2026) -
Differentiate-and-Inject: Enhancing VLAs via Functional Differentiation Induced by In-Parameter Structural Reasoning
von: Hou, Jingyi, et al.
Veröffentlicht: (2026) -
ARCADE: Scalable Demonstration Collection and Generation via Augmented Reality for Imitation Learning
von: Yang, Yue, et al.
Veröffentlicht: (2024)