VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Chaoyang, Bao, Wenrui, Gao, Sicheng, Xu, Bingxin, Tian, Yu, Rawat, Yogesh S., Ge, Yunhao, Shang, Yuzhang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SilentDrift: Exploiting Action Chunking for Stealthy Backdoor Attacks on Vision-Language-Action Models
von: Xu, Bingxin, et al.
Veröffentlicht: (2026)
von: Xu, Bingxin, et al.
Veröffentlicht: (2026)
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
von: Bai, Shuanghao, et al.
Veröffentlicht: (2026)
von: Bai, Shuanghao, et al.
Veröffentlicht: (2026)
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
von: Yin, Cheng, et al.
Veröffentlicht: (2025)
von: Yin, Cheng, et al.
Veröffentlicht: (2025)
Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
von: Peng, Zhenghao "Mark", et al.
Veröffentlicht: (2025)
von: Peng, Zhenghao "Mark", et al.
Veröffentlicht: (2025)
ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models
von: Li, Ye, et al.
Veröffentlicht: (2026)
von: Li, Ye, et al.
Veröffentlicht: (2026)
OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning
von: Lin, Fanqi, et al.
Veröffentlicht: (2025)
von: Lin, Fanqi, et al.
Veröffentlicht: (2025)
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
von: Sun, Jianli, et al.
Veröffentlicht: (2026)
von: Sun, Jianli, et al.
Veröffentlicht: (2026)
SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios
von: Gao, Tian, et al.
Veröffentlicht: (2026)
von: Gao, Tian, et al.
Veröffentlicht: (2026)
DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
von: Zheng, Zihao, et al.
Veröffentlicht: (2026)
TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments
von: Huang, Zhiyu, et al.
Veröffentlicht: (2026)
von: Huang, Zhiyu, et al.
Veröffentlicht: (2026)
RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models
von: Koo, Jiyeon, et al.
Veröffentlicht: (2025)
von: Koo, Jiyeon, et al.
Veröffentlicht: (2025)
RationalVLA: A Rational Vision-Language-Action Model with Dual System
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
von: Xu, Xiaoxu, et al.
Veröffentlicht: (2026)
von: Xu, Xiaoxu, et al.
Veröffentlicht: (2026)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
von: Zhong, Linqing, et al.
Veröffentlicht: (2026)
TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
von: Zhang, Zongzheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zongzheng, et al.
Veröffentlicht: (2025)
VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
von: Ye, Angen, et al.
Veröffentlicht: (2025)
von: Ye, Angen, et al.
Veröffentlicht: (2025)
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
von: Li, Boyu, et al.
Veröffentlicht: (2026)
von: Li, Boyu, et al.
Veröffentlicht: (2026)
FlowVLA: Visual Chain of Thought-based Motion Reasoning for Vision-Language-Action Models
von: Zhong, Zhide, et al.
Veröffentlicht: (2025)
von: Zhong, Zhide, et al.
Veröffentlicht: (2025)
VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
von: Guo, Wenkai, et al.
Veröffentlicht: (2025)
von: Guo, Wenkai, et al.
Veröffentlicht: (2025)
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
von: Deng, Shengliang, et al.
Veröffentlicht: (2025)
von: Deng, Shengliang, et al.
Veröffentlicht: (2025)
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
von: Hu, Xintong, et al.
Veröffentlicht: (2026)
von: Hu, Xintong, et al.
Veröffentlicht: (2026)
ProGAL-VLA: Grounded Alignment through Prospective Reasoning in Vision-Language-Action Models
von: Darabi, Nastaran, et al.
Veröffentlicht: (2026)
von: Darabi, Nastaran, et al.
Veröffentlicht: (2026)
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
von: Wang, Yihao, et al.
Veröffentlicht: (2025)
von: Wang, Yihao, et al.
Veröffentlicht: (2025)
ST4VLA: Spatially Guided Training for Vision-Language-Action Models
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
UnderwaterVLA: Dual-brain Vision-Language-Action architecture for Autonomous Underwater Navigation
von: Wang, Zhangyuan, et al.
Veröffentlicht: (2025)
von: Wang, Zhangyuan, et al.
Veröffentlicht: (2025)
StyleVLA: Driving Style-Aware Vision Language Action Model for Autonomous Driving
von: Gao, Yuan, et al.
Veröffentlicht: (2026)
von: Gao, Yuan, et al.
Veröffentlicht: (2026)
EdgeVLA: Efficient Vision-Language-Action Models
von: Budzianowski, Paweł, et al.
Veröffentlicht: (2025)
von: Budzianowski, Paweł, et al.
Veröffentlicht: (2025)
SELF-VLA: A Skill Enhanced Agentic Vision-Language-Action Framework for Contact-Rich Disassembly
von: Liu, Chang, et al.
Veröffentlicht: (2026)
von: Liu, Chang, et al.
Veröffentlicht: (2026)
Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
SwitchVLA: Execution-Aware Task Switching for Vision-Language-Action Models
von: Li, Meng, et al.
Veröffentlicht: (2025)
von: Li, Meng, et al.
Veröffentlicht: (2025)
Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation
von: Wei, Xiangyi, et al.
Veröffentlicht: (2025)
von: Wei, Xiangyi, et al.
Veröffentlicht: (2025)
HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
von: Lin, Minghui, et al.
Veröffentlicht: (2025)
von: Lin, Minghui, et al.
Veröffentlicht: (2025)
SmoothVLA: Aligning Vision-Language-Action Models with Physical Constraints via Intrinsic Smoothness Optimization
von: Li, Jiashun, et al.
Veröffentlicht: (2026)
von: Li, Jiashun, et al.
Veröffentlicht: (2026)
DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation
von: Peng, Xiongfeng, et al.
Veröffentlicht: (2026)
von: Peng, Xiongfeng, et al.
Veröffentlicht: (2026)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
von: Xu, Zonghuan, et al.
Veröffentlicht: (2025)
von: Xu, Zonghuan, et al.
Veröffentlicht: (2025)
AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
von: Jiang, Yuhua, et al.
Veröffentlicht: (2025)
von: Jiang, Yuhua, et al.
Veröffentlicht: (2025)
PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models
von: Guo, Xinyu, et al.
Veröffentlicht: (2026)
von: Guo, Xinyu, et al.
Veröffentlicht: (2026)
FocusVLA: Focused Visual Utilization for Vision-Language-Action Models
von: Zhang, Yichi, et al.
Veröffentlicht: (2026)
von: Zhang, Yichi, et al.
Veröffentlicht: (2026)
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
von: Wang, Zixuan, et al.
Veröffentlicht: (2026)
von: Wang, Zixuan, et al.
Veröffentlicht: (2026)
RedVLA: Physical Red Teaming for Vision-Language-Action Models
von: Zhang, Yuhao, et al.
Veröffentlicht: (2026)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SilentDrift: Exploiting Action Chunking for Stealthy Backdoor Attacks on Vision-Language-Action Models
von: Xu, Bingxin, et al.
Veröffentlicht: (2026) -
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
von: Bai, Shuanghao, et al.
Veröffentlicht: (2026) -
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
von: Yin, Cheng, et al.
Veröffentlicht: (2025) -
Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning
von: Peng, Zhenghao "Mark", et al.
Veröffentlicht: (2025) -
ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models
von: Li, Ye, et al.
Veröffentlicht: (2026)