KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zheng, Zihao, Mao, Zhihao, Li, Maoliang, Chen, Jiayu, Sun, Xinhao, Zhang, Zhaobo, Cao, Donggang, Mei, Hong, Chen, Xiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
FreqCache: Accelerating Embodied VLN Models with Adaptive Frequency-Guided Token Caching
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
EaqVLA: Encoding-aligned Quantization for Vision-Language-Action Models
di: Jiang, Feng, et al.
Pubblicazione: (2025)
di: Jiang, Feng, et al.
Pubblicazione: (2025)
Anticipation-VLA: Solving Long-Horizon Embodied Tasks via Anticipation-based Subgoal Generation
di: Zhang, Zhilong, et al.
Pubblicazione: (2026)
di: Zhang, Zhilong, et al.
Pubblicazione: (2026)
Vidarc: Embodied Video Diffusion Model for Closed-loop Control
di: Feng, Yao, et al.
Pubblicazione: (2025)
di: Feng, Yao, et al.
Pubblicazione: (2025)
SimVLA: A Simple VLA Baseline for Robotic Manipulation
di: Luo, Yuankai, et al.
Pubblicazione: (2026)
di: Luo, Yuankai, et al.
Pubblicazione: (2026)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
di: Zheng, Zihao, et al.
Pubblicazione: (2026)
AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps
di: Fan, Liaoyuan, et al.
Pubblicazione: (2026)
di: Fan, Liaoyuan, et al.
Pubblicazione: (2026)
STaRR: Spatial-Temporal Token-Dynamics-Aware Responsive Remasking for Diffusion Language Models
di: Sun, Xinhao, et al.
Pubblicazione: (2025)
di: Sun, Xinhao, et al.
Pubblicazione: (2025)
Can VLA Models Learn from Real-World Data Continually without Forgetting?
di: Zhu, Jiarun, et al.
Pubblicazione: (2026)
di: Zhu, Jiarun, et al.
Pubblicazione: (2026)
DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset Adaptation
di: Zheng, Zihao, et al.
Pubblicazione: (2025)
di: Zheng, Zihao, et al.
Pubblicazione: (2025)
TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
di: Liu, Jiahang, et al.
Pubblicazione: (2025)
di: Liu, Jiahang, et al.
Pubblicazione: (2025)
AtomVLA: Scalable Post-Training for Robotic Manipulation via Predictive Latent World Models
di: Sun, Xiaoquan, et al.
Pubblicazione: (2026)
di: Sun, Xiaoquan, et al.
Pubblicazione: (2026)
Embodied Red Teaming for Auditing Robotic Foundation Models
di: Karnik, Sathwik, et al.
Pubblicazione: (2024)
di: Karnik, Sathwik, et al.
Pubblicazione: (2024)
AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge
di: Hirose, Noriaki, et al.
Pubblicazione: (2026)
di: Hirose, Noriaki, et al.
Pubblicazione: (2026)
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
di: Liang, Zhixuan, et al.
Pubblicazione: (2025)
di: Liang, Zhixuan, et al.
Pubblicazione: (2025)
VLA-Touch: Enhancing Vision-Language-Action Models with Dual-Level Tactile Feedback
di: Bi, Jianxin, et al.
Pubblicazione: (2025)
di: Bi, Jianxin, et al.
Pubblicazione: (2025)
From Inference Efficiency to Embodied Efficiency: Revisiting Efficiency Metrics for Vision-Language-Action Models
di: Li, Zhuofan, et al.
Pubblicazione: (2026)
di: Li, Zhuofan, et al.
Pubblicazione: (2026)
Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
di: Yuan, Yifu, et al.
Pubblicazione: (2025)
di: Yuan, Yifu, et al.
Pubblicazione: (2025)
NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks
di: Luo, Zhihao, et al.
Pubblicazione: (2025)
di: Luo, Zhihao, et al.
Pubblicazione: (2025)
Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic Manipulation
di: Lv, Qi, et al.
Pubblicazione: (2025)
di: Lv, Qi, et al.
Pubblicazione: (2025)
MetaVLA: Unified Meta Co-training For Efficient Embodied Adaption
di: Li, Chen, et al.
Pubblicazione: (2025)
di: Li, Chen, et al.
Pubblicazione: (2025)
D-CAT: Decoupled Cross-Attention Transfer between Sensor Modalities for Unimodal Inference
di: Daher, Leen, et al.
Pubblicazione: (2025)
di: Daher, Leen, et al.
Pubblicazione: (2025)
Self-Improving Embodied Foundation Models
di: Ghasemipour, Seyed Kamyar Seyed, et al.
Pubblicazione: (2025)
di: Ghasemipour, Seyed Kamyar Seyed, et al.
Pubblicazione: (2025)
OpenVLA: An Open-Source Vision-Language-Action Model
di: Kim, Moo Jin, et al.
Pubblicazione: (2024)
di: Kim, Moo Jin, et al.
Pubblicazione: (2024)
CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
Shifting Uncertainty to Critical Moments: Towards Reliable Uncertainty Quantification for VLA Model
di: Tang, Yanchuan, et al.
Pubblicazione: (2026)
di: Tang, Yanchuan, et al.
Pubblicazione: (2026)
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
di: Feng, Yao, et al.
Pubblicazione: (2025)
di: Feng, Yao, et al.
Pubblicazione: (2025)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
di: Xu, Zonghuan, et al.
Pubblicazione: (2025)
di: Xu, Zonghuan, et al.
Pubblicazione: (2025)
HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
di: Xiong, Zheng, et al.
Pubblicazione: (2025)
di: Xiong, Zheng, et al.
Pubblicazione: (2025)
UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models
di: Zhang, Qiyao, et al.
Pubblicazione: (2026)
di: Zhang, Qiyao, et al.
Pubblicazione: (2026)
Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
di: Jabbour, Jason, et al.
Pubblicazione: (2025)
di: Jabbour, Jason, et al.
Pubblicazione: (2025)
Robotic Control via Embodied Chain-of-Thought Reasoning
di: Zawalski, Michał, et al.
Pubblicazione: (2024)
di: Zawalski, Michał, et al.
Pubblicazione: (2024)
Model Adaptation for Time Constrained Embodied Control
di: Song, Jaehyun, et al.
Pubblicazione: (2024)
di: Song, Jaehyun, et al.
Pubblicazione: (2024)
VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model
di: Li, Wenhao, et al.
Pubblicazione: (2026)
di: Li, Wenhao, et al.
Pubblicazione: (2026)
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
di: Li, Haozhan, et al.
Pubblicazione: (2025)
di: Li, Haozhan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness
di: Zheng, Zihao, et al.
Pubblicazione: (2026) -
DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models
di: Zheng, Zihao, et al.
Pubblicazione: (2026) -
VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness
di: Zheng, Zihao, et al.
Pubblicazione: (2026) -
RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models
di: Zheng, Zihao, et al.
Pubblicazione: (2026) -
RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models
di: Zheng, Zihao, et al.
Pubblicazione: (2026)