HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Zihao, Mao, Zhihao, Tian, Sicheng, Li, Maoliang, Chen, Jiayu, Sun, Xinhao, Zhang, Zhaobo, Liu, Xuanzhe, Cao, Donggang, Mei, Hong, Chen, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
STaRR: Spatial-Temporal Token-Dynamics-Aware Responsive Remasking for Diffusion Language Models
by: Sun, Xinhao, et al.
Published: (2025)
by: Sun, Xinhao, et al.
Published: (2025)
EaqVLA: Encoding-aligned Quantization for Vision-Language-Action Models
by: Jiang, Feng, et al.
Published: (2025)
by: Jiang, Feng, et al.
Published: (2025)
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
by: Li, Maoliang, et al.
Published: (2026)
by: Li, Maoliang, et al.
Published: (2026)
StarSD: One-for-Many Speculative Decoding
by: He, Junhao, et al.
Published: (2026)
by: He, Junhao, et al.
Published: (2026)
DReSD: Dense Retrieval for Speculative Decoding
by: Gritta, Milan, et al.
Published: (2025)
by: Gritta, Milan, et al.
Published: (2025)
ToProVAR: Efficient Visual Autoregressive Modeling via Tri-Dimensional Entropy-Aware Semantic Analysis and Sparsity Optimization
by: Chen, Jiayu, et al.
Published: (2026)
by: Chen, Jiayu, et al.
Published: (2026)
FedHQ: Hybrid Runtime Quantization for Federated Learning
by: Zheng, Zihao, et al.
Published: (2025)
by: Zheng, Zihao, et al.
Published: (2025)
DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset Adaptation
by: Zheng, Zihao, et al.
Published: (2025)
by: Zheng, Zihao, et al.
Published: (2025)
Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation
by: Chen, Jiayu, et al.
Published: (2026)
by: Chen, Jiayu, et al.
Published: (2026)
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
by: Kang, Jialiang, et al.
Published: (2025)
by: Kang, Jialiang, et al.
Published: (2025)
TPP-SD: Accelerating Transformer Point Process Sampling with Speculative Decoding
by: Gong, Shukai, et al.
Published: (2025)
by: Gong, Shukai, et al.
Published: (2025)
AdaSD: Adaptive Speculative Decoding for Efficient Language Model Inference
by: Lu, Kuan-Wei, et al.
Published: (2025)
by: Lu, Kuan-Wei, et al.
Published: (2025)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
by: Li, Zikun, et al.
Published: (2025)
by: Li, Zikun, et al.
Published: (2025)
Computational geometry in Heisenberg group Heis^3
by: Marenich, Andrey
Published: (2002)
by: Marenich, Andrey
Published: (2002)
Traversal Verification for Speculative Tree Decoding
by: Weng, Yepeng, et al.
Published: (2025)
by: Weng, Yepeng, et al.
Published: (2025)
MIRAGE: Runtime Scheduling for Multi-Vector Image Retrieval with Hierarchical Decomposition
by: Li, Maoliang, et al.
Published: (2025)
by: Li, Maoliang, et al.
Published: (2025)
Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
by: Chen, Zhuoming, et al.
Published: (2024)
by: Chen, Zhuoming, et al.
Published: (2024)
DiP-SD: Distributed Pipelined Speculative Decoding for Efficient LLM Inference at the Edge
by: Xu, Yaodan, et al.
Published: (2026)
by: Xu, Yaodan, et al.
Published: (2026)
EMS-SD: Efficient Multi-sample Speculative Decoding for Accelerating Large Language Models
by: Ni, Yunsheng, et al.
Published: (2024)
by: Ni, Yunsheng, et al.
Published: (2024)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
by: Han, Yunhe, et al.
Published: (2026)
by: Han, Yunhe, et al.
Published: (2026)
Component-Aware Self-Speculative Decoding in Hybrid Language Models
by: Borobia, Hector, et al.
Published: (2026)
by: Borobia, Hector, et al.
Published: (2026)
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
by: Oliaro, Gabriele, et al.
Published: (2024)
by: Oliaro, Gabriele, et al.
Published: (2024)
Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding
by: Sun, Shuoyang, et al.
Published: (2026)
by: Sun, Shuoyang, et al.
Published: (2026)
Speculative Safety-Aware Decoding
by: Wang, Xuekang, et al.
Published: (2025)
by: Wang, Xuekang, et al.
Published: (2025)
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
by: Wang, Songsheng, et al.
Published: (2025)
by: Wang, Songsheng, et al.
Published: (2025)
SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
by: Wang, Hanzhen, et al.
Published: (2025)
by: Wang, Hanzhen, et al.
Published: (2025)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
Scaling Laws for Speculative Decoding
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
CASCADE: Context-Aware Relaxation for Speculative Image Decoding
by: Yildirim, Selin, et al.
Published: (2026)
by: Yildirim, Selin, et al.
Published: (2026)
SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
by: Svirschevski, Ruslan, et al.
Published: (2024)
by: Svirschevski, Ruslan, et al.
Published: (2024)
RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding
by: Hoshino, Yuichiro, et al.
Published: (2025)
by: Hoshino, Yuichiro, et al.
Published: (2025)
Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding
by: Su, Xin, et al.
Published: (2026)
by: Su, Xin, et al.
Published: (2026)
Speculative Speculative Decoding
by: Kumar, Tanishq, et al.
Published: (2026)
by: Kumar, Tanishq, et al.
Published: (2026)
TAPS: Target-Aware Prefix Tree Selection for Diffusion-Drafted Speculative Decoding
by: Wang, Zhuoyu, et al.
Published: (2026)
by: Wang, Zhuoyu, et al.
Published: (2026)
Similar Items
-
KERV: Kinematic-Rectified Speculative Decoding for Embodied VLA Models
by: Zheng, Zihao, et al.
Published: (2026) -
DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models
by: Zheng, Zihao, et al.
Published: (2026) -
VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness
by: Zheng, Zihao, et al.
Published: (2026) -
RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models
by: Zheng, Zihao, et al.
Published: (2026) -
RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models
by: Zheng, Zihao, et al.
Published: (2026)