Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Hanxin, Xu, Mingshuo, Dhafer, Abdulqader, Yue, Shigang, Dong, Hongbiao, Hao, Zhou Daniel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Generative System for Robot-to-Human Handovers: from Intent Inference to Spatial Configuration Imagery
di: Zhang, Hanxin, et al.
Pubblicazione: (2025)
di: Zhang, Hanxin, et al.
Pubblicazione: (2025)
Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models
di: Xu, Haiweng, et al.
Pubblicazione: (2026)
di: Xu, Haiweng, et al.
Pubblicazione: (2026)
Survey of Vision-Language-Action Models for Embodied Manipulation
di: Li, Haoran, et al.
Pubblicazione: (2025)
di: Li, Haoran, et al.
Pubblicazione: (2025)
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
di: Li, Boyu, et al.
Pubblicazione: (2026)
di: Li, Boyu, et al.
Pubblicazione: (2026)
DM0: An Embodied-Native Vision-Language-Action Model towards Physical AI
di: Yu, En, et al.
Pubblicazione: (2026)
di: Yu, En, et al.
Pubblicazione: (2026)
Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents
di: Yang, Zhejian, et al.
Pubblicazione: (2025)
di: Yang, Zhejian, et al.
Pubblicazione: (2025)
HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning
di: Shou, Quanxin, et al.
Pubblicazione: (2026)
di: Shou, Quanxin, et al.
Pubblicazione: (2026)
Embodied Learning of Reward for Musculoskeletal Control with Vision Language Models
di: Soedarmadji, Saraswati, et al.
Pubblicazione: (2025)
di: Soedarmadji, Saraswati, et al.
Pubblicazione: (2025)
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
di: Sun, Jingwen, et al.
Pubblicazione: (2026)
di: Sun, Jingwen, et al.
Pubblicazione: (2026)
F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
di: Lv, Qi, et al.
Pubblicazione: (2025)
di: Lv, Qi, et al.
Pubblicazione: (2025)
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
di: Torne, Marcel, et al.
Pubblicazione: (2026)
di: Torne, Marcel, et al.
Pubblicazione: (2026)
Embodied Scene Understanding for Vision Language Models via MetaVQA
di: Wang, Weizhen, et al.
Pubblicazione: (2025)
di: Wang, Weizhen, et al.
Pubblicazione: (2025)
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
di: Chen, William, et al.
Pubblicazione: (2026)
di: Chen, William, et al.
Pubblicazione: (2026)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
di: Ling, Yiran, et al.
Pubblicazione: (2026)
di: Ling, Yiran, et al.
Pubblicazione: (2026)
Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
di: Guan, Weifan, et al.
Pubblicazione: (2025)
di: Guan, Weifan, et al.
Pubblicazione: (2025)
TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
di: Zhang, Zongzheng, et al.
Pubblicazione: (2025)
di: Zhang, Zongzheng, et al.
Pubblicazione: (2025)
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models
di: Zhang, Jiyao, et al.
Pubblicazione: (2026)
di: Zhang, Jiyao, et al.
Pubblicazione: (2026)
A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter
di: Xu, Kechun, et al.
Pubblicazione: (2023)
di: Xu, Kechun, et al.
Pubblicazione: (2023)
Toward Embodiment Equivariant Vision-Language-Action Policy
di: Chen, Anzhe, et al.
Pubblicazione: (2025)
di: Chen, Anzhe, et al.
Pubblicazione: (2025)
Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation
di: Ding, Hongyu, et al.
Pubblicazione: (2026)
di: Ding, Hongyu, et al.
Pubblicazione: (2026)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
di: Li, Xiaoqi, et al.
Pubblicazione: (2025)
di: Li, Xiaoqi, et al.
Pubblicazione: (2025)
DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
di: Yang, Yurou, et al.
Pubblicazione: (2026)
di: Yang, Yurou, et al.
Pubblicazione: (2026)
UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models
di: Zhang, Qiyao, et al.
Pubblicazione: (2026)
di: Zhang, Qiyao, et al.
Pubblicazione: (2026)
InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning
di: Zhang, Ji, et al.
Pubblicazione: (2025)
di: Zhang, Ji, et al.
Pubblicazione: (2025)
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models
di: Li, Xiaoqi, et al.
Pubblicazione: (2026)
di: Li, Xiaoqi, et al.
Pubblicazione: (2026)
RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI
di: Tai, Cong, et al.
Pubblicazione: (2025)
di: Tai, Cong, et al.
Pubblicazione: (2025)
HMR-1: Hierarchical Massage Robot with Vision-Language-Model for Embodied Healthcare
di: Xu, Rongtao, et al.
Pubblicazione: (2026)
di: Xu, Rongtao, et al.
Pubblicazione: (2026)
Stable Language Guidance for Vision-Language-Action Models
di: Zhan, Zhihao, et al.
Pubblicazione: (2026)
di: Zhan, Zhihao, et al.
Pubblicazione: (2026)
Action Hallucination in Generative Vision-Language-Action Models
di: Soh, Harold, et al.
Pubblicazione: (2026)
di: Soh, Harold, et al.
Pubblicazione: (2026)
HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
From Inference Efficiency to Embodied Efficiency: Revisiting Efficiency Metrics for Vision-Language-Action Models
di: Li, Zhuofan, et al.
Pubblicazione: (2026)
di: Li, Zhuofan, et al.
Pubblicazione: (2026)
VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation
di: Zhang, Chaofan, et al.
Pubblicazione: (2025)
di: Zhang, Chaofan, et al.
Pubblicazione: (2025)
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
di: Liu, Zhuoyang, et al.
Pubblicazione: (2025)
di: Liu, Zhuoyang, et al.
Pubblicazione: (2025)
FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
di: Xu, Xiaoxu, et al.
Pubblicazione: (2026)
di: Xu, Xiaoxu, et al.
Pubblicazione: (2026)
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models
di: Lei, Zixing, et al.
Pubblicazione: (2026)
di: Lei, Zixing, et al.
Pubblicazione: (2026)
AnchorVLA4D: an Anchor-Based Spatial-Temporal Vision-Language-Action Model for Robotic Manipulation
di: Zhu, Juan, et al.
Pubblicazione: (2026)
di: Zhu, Juan, et al.
Pubblicazione: (2026)
Grounding Sim-to-Real Generalization in Dexterous Manipulation: An Empirical Study with Vision-Language-Action Models
di: Jin, Ruixing, et al.
Pubblicazione: (2026)
di: Jin, Ruixing, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A Generative System for Robot-to-Human Handovers: from Intent Inference to Spatial Configuration Imagery
di: Zhang, Hanxin, et al.
Pubblicazione: (2025) -
Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models
di: Xu, Haiweng, et al.
Pubblicazione: (2026) -
Survey of Vision-Language-Action Models for Embodied Manipulation
di: Li, Haoran, et al.
Pubblicazione: (2025) -
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
di: Li, Boyu, et al.
Pubblicazione: (2026) -
DM0: An Embodied-Native Vision-Language-Action Model towards Physical AI
di: Yu, En, et al.
Pubblicazione: (2026)