Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Haiweng, Zheng, Sipeng, Luo, Hao, Zhang, Wanpeng, Xi, Ziheng, Lu, Zongqing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908979942129664
author Xu, Haiweng
Zheng, Sipeng
Luo, Hao
Zhang, Wanpeng
Xi, Ziheng
Lu, Zongqing
author_facet Xu, Haiweng
Zheng, Sipeng
Luo, Hao
Zhang, Wanpeng
Xi, Ziheng
Lu, Zongqing
contents Recent Vision-Language-Action (VLA) models report impressive success rates on standard robotic benchmarks, fueling optimism about general-purpose physical intelligence. However, recent evidence suggests a systematic misalignment between standard benchmark success and true embodied reasoning, raising the question of whether these high scores reflect genuine cognitive capability. To address this gap, we introduce BeTTER, a diagnostic Benchmark for Testing True Embodied Reasoning in robotic policies. BeTTER applies targeted causal interventions (e.g., spatial layout shifts, temporal extrapolation) while enforcing kinematic isolation to explicitly decouple high-level reasoning failures from low-level execution limits. Through systematic evaluation, we reveal that state-of-the-art VLAs catastrophically fail in dynamic scenarios, exhibiting severe lexical-kinematic shortcuts, behavioral inertia, and semantic feature collapse. Crucially, our mechanistic analysis traces these symptoms to fundamental architectural bottlenecks - such as capacity compression and myopic downsampling - which systematically degrade the model's foundational semantic representation. We demonstrate that highly static evaluation protocols effectively mask this degradation by allowing optimization to overfit to sensorimotor priors. Supported by real-world robotic validation, our findings confirm that this representational breakdown is not a simulation artifact, highlighting the critical need for future VLA paradigms to resolve the structural tension between high-frequency control and high-level reasoning.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18000
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models
Xu, Haiweng
Zheng, Sipeng
Luo, Hao
Zhang, Wanpeng
Xi, Ziheng
Lu, Zongqing
Robotics
Recent Vision-Language-Action (VLA) models report impressive success rates on standard robotic benchmarks, fueling optimism about general-purpose physical intelligence. However, recent evidence suggests a systematic misalignment between standard benchmark success and true embodied reasoning, raising the question of whether these high scores reflect genuine cognitive capability. To address this gap, we introduce BeTTER, a diagnostic Benchmark for Testing True Embodied Reasoning in robotic policies. BeTTER applies targeted causal interventions (e.g., spatial layout shifts, temporal extrapolation) while enforcing kinematic isolation to explicitly decouple high-level reasoning failures from low-level execution limits. Through systematic evaluation, we reveal that state-of-the-art VLAs catastrophically fail in dynamic scenarios, exhibiting severe lexical-kinematic shortcuts, behavioral inertia, and semantic feature collapse. Crucially, our mechanistic analysis traces these symptoms to fundamental architectural bottlenecks - such as capacity compression and myopic downsampling - which systematically degrade the model's foundational semantic representation. We demonstrate that highly static evaluation protocols effectively mask this degradation by allowing optimization to overfit to sensorimotor priors. Supported by real-world robotic validation, our findings confirm that this representational breakdown is not a simulation artifact, highlighting the critical need for future VLA paradigms to resolve the structural tension between high-frequency control and high-level reasoning.
title Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models
topic Robotics
url https://arxiv.org/abs/2604.18000