Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Bai, Shuanghao, Wang, Dakai, Chi, Cheng, Zhou, Wanqi, Lyu, Jing, Zhao, Xiaoguang, Wang, Pengwei, Wang, Zhongyuan, Xing, Lei, Zhang, Shanghang, Chen, Badong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
di: Fu, Yankai, et al.
Pubblicazione: (2025)
di: Fu, Yankai, et al.
Pubblicazione: (2025)
VCoT-Grasp: Grasp Foundation Models with Visual Chain-of-Thought Reasoning for Language-driven Grasp Generation
di: Zhang, Haoran, et al.
Pubblicazione: (2025)
di: Zhang, Haoran, et al.
Pubblicazione: (2025)
Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation
di: Tan, Huajie, et al.
Pubblicazione: (2026)
di: Tan, Huajie, et al.
Pubblicazione: (2026)
Assistance Without Interruption: A Benchmark and LLM-based Framework for Non-Intrusive Human-Robot Assistance
di: Zhang, Yuedi, et al.
Pubblicazione: (2026)
di: Zhang, Yuedi, et al.
Pubblicazione: (2026)
Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation
di: Fan, Yiguo, et al.
Pubblicazione: (2025)
di: Fan, Yiguo, et al.
Pubblicazione: (2025)
VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
di: Zhao, Wei, et al.
Pubblicazione: (2025)
di: Zhao, Wei, et al.
Pubblicazione: (2025)
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
di: Han, Yi, et al.
Pubblicazione: (2025)
di: Han, Yi, et al.
Pubblicazione: (2025)
BlockVLA: Accelerating Autoregressive VLA via Block Diffusion Finetuning
di: Wang, Ruiheng, et al.
Pubblicazione: (2026)
di: Wang, Ruiheng, et al.
Pubblicazione: (2026)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
OmniSAT: Compact Action Token, Faster Auto Regression
di: Lyu, Huaihai, et al.
Pubblicazione: (2025)
di: Lyu, Huaihai, et al.
Pubblicazione: (2025)
MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation
di: Zhang, Lingfeng, et al.
Pubblicazione: (2025)
di: Zhang, Lingfeng, et al.
Pubblicazione: (2025)
Stable Language Guidance for Vision-Language-Action Models
di: Zhan, Zhihao, et al.
Pubblicazione: (2026)
di: Zhan, Zhihao, et al.
Pubblicazione: (2026)
Revisiting the Adversarial Robustness of Vision Language Models: a Multimodal Perspective
di: Zhou, Wanqi, et al.
Pubblicazione: (2024)
di: Zhou, Wanqi, et al.
Pubblicazione: (2024)
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
di: Zhou, Enshen, et al.
Pubblicazione: (2025)
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
di: Li, Zhe, et al.
Pubblicazione: (2025)
di: Li, Zhe, et al.
Pubblicazione: (2025)
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
di: Deng, Shengliang, et al.
Pubblicazione: (2025)
di: Deng, Shengliang, et al.
Pubblicazione: (2025)
ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
di: Dai, Yuntao, et al.
Pubblicazione: (2025)
di: Dai, Yuntao, et al.
Pubblicazione: (2025)
HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
di: Bai, Shuanghao, et al.
Pubblicazione: (2026)
PIGEON: VLM-Driven Object Navigation via Points of Interest Selection
di: Peng, Cheng, et al.
Pubblicazione: (2025)
di: Peng, Cheng, et al.
Pubblicazione: (2025)
Mean-Flow based One-Step Vision-Language-Action
di: Chen, Yang, et al.
Pubblicazione: (2026)
di: Chen, Yang, et al.
Pubblicazione: (2026)
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
di: Bai, Shuanghao, et al.
Pubblicazione: (2025)
Adaptive Action Chunking at Inference-time for Vision-Language-Action Models
di: Liang, Yuanchang, et al.
Pubblicazione: (2026)
di: Liang, Yuanchang, et al.
Pubblicazione: (2026)
AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter
di: Tang, Yingbo, et al.
Pubblicazione: (2025)
di: Tang, Yingbo, et al.
Pubblicazione: (2025)
World-Value-Action Model: Implicit Planning for Vision-Language-Action Systems
di: Li, Runze, et al.
Pubblicazione: (2026)
di: Li, Runze, et al.
Pubblicazione: (2026)
From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
di: Li, Zhe, et al.
Pubblicazione: (2025)
di: Li, Zhe, et al.
Pubblicazione: (2025)
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
A Survey on Vision-Language-Action Models: An Action Tokenization Perspective
di: Zhong, Yifan, et al.
Pubblicazione: (2025)
di: Zhong, Yifan, et al.
Pubblicazione: (2025)
LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
di: Li, Zuolei, et al.
Pubblicazione: (2025)
di: Li, Zuolei, et al.
Pubblicazione: (2025)
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
di: Song, Wenxuan, et al.
Pubblicazione: (2026)
di: Song, Wenxuan, et al.
Pubblicazione: (2026)
EndoVLA: Dual-Phase Vision-Language-Action Model for Autonomous Tracking in Endoscopy
di: Ng, Chi Kit, et al.
Pubblicazione: (2025)
di: Ng, Chi Kit, et al.
Pubblicazione: (2025)
VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic Manipulation
di: Wang, Zhijie, et al.
Pubblicazione: (2024)
di: Wang, Zhijie, et al.
Pubblicazione: (2024)
Toward Embodiment Equivariant Vision-Language-Action Policy
di: Chen, Anzhe, et al.
Pubblicazione: (2025)
di: Chen, Anzhe, et al.
Pubblicazione: (2025)
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation
di: Yu, Wenda, et al.
Pubblicazione: (2026)
di: Yu, Wenda, et al.
Pubblicazione: (2026)
Unified Vision-Language-Action Model
di: Wang, Yuqi, et al.
Pubblicazione: (2025)
di: Wang, Yuqi, et al.
Pubblicazione: (2025)
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
di: Zhang, Jiazhao, et al.
Pubblicazione: (2024)
di: Zhang, Jiazhao, et al.
Pubblicazione: (2024)
ProbeFlow: Training-Free Adaptive Flow Matching for Vision-Language-Action Models
di: Fang, Zhou, et al.
Pubblicazione: (2026)
di: Fang, Zhou, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
di: Bai, Shuanghao, et al.
Pubblicazione: (2026) -
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
di: Liu, Mengzhen, et al.
Pubblicazione: (2026) -
Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation
di: Bai, Shuanghao, et al.
Pubblicazione: (2025) -
METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
di: Fu, Yankai, et al.
Pubblicazione: (2025) -
VCoT-Grasp: Grasp Foundation Models with Visual Chain-of-Thought Reasoning for Language-driven Grasp Generation
di: Zhang, Haoran, et al.
Pubblicazione: (2025)