HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Zhinan, Wang, Peisong, Qiu, Shuang, Cheng, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HiSpec: Hierarchical Speculative Decoding for LLMs
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative Drafter
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
Recurrent Drafter for Fast Speculative Decoding in Large Language Models
von: Cheng, Yunfei, et al.
Veröffentlicht: (2024)
von: Cheng, Yunfei, et al.
Veröffentlicht: (2024)
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
von: Wang, Songsheng, et al.
Veröffentlicht: (2025)
von: Wang, Songsheng, et al.
Veröffentlicht: (2025)
Mamba Drafters for Speculative Decoding
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
SparVAR: Exploring Sparsity in Visual AutoRegressive Modeling for Training-Free Acceleration
von: Li, Zekun, et al.
Veröffentlicht: (2026)
von: Li, Zekun, et al.
Veröffentlicht: (2026)
Speculative Safety-Aware Decoding
von: Wang, Xuekang, et al.
Veröffentlicht: (2025)
von: Wang, Xuekang, et al.
Veröffentlicht: (2025)
FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
von: Bajpai, Divya Jyoti, et al.
Veröffentlicht: (2025)
On Speculative Decoding for Multimodal Large Language Models
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
Steering Pretrained Drafters during Speculative Decoding
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2025)
Confidence-Modulated Speculative Decoding for Large Language Models
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2025)
von: Li, Jinze, et al.
Veröffentlicht: (2025)
Token-Driven GammaTune: Adaptive Calibration for Enhanced Speculative Decoding
von: Gautam, Aayush, et al.
Veröffentlicht: (2025)
von: Gautam, Aayush, et al.
Veröffentlicht: (2025)
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
von: Shen, Yuhao, et al.
Veröffentlicht: (2026)
von: Shen, Yuhao, et al.
Veröffentlicht: (2026)
Traversal Verification for Speculative Tree Decoding
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
QSpec: Speculative Decoding with Complementary Quantization Schemes
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
Fast Large Language Model Collaborative Decoding via Speculation
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
STree: Speculative Tree Decoding for Hybrid State-Space Models
von: Wu, Yangchao, et al.
Veröffentlicht: (2025)
von: Wu, Yangchao, et al.
Veröffentlicht: (2025)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
von: Liu, Hongyi, et al.
Veröffentlicht: (2025)
von: Liu, Hongyi, et al.
Veröffentlicht: (2025)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
Accelerating PayPal's Commerce Agent with Speculative Decoding: An Empirical Study on EAGLE3 with Fine-Tuned Nemotron Models
von: Qin, Ally, et al.
Veröffentlicht: (2026)
von: Qin, Ally, et al.
Veröffentlicht: (2026)
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
von: Shi, Lucy Xiaoyang, et al.
Veröffentlicht: (2025)
von: Shi, Lucy Xiaoyang, et al.
Veröffentlicht: (2025)
Ban&Pick: Ehancing Performance and Efficiency of MoE-LLMs via Smarter Routing
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
Coupling without Communication and Drafter-Invariant Speculative Decoding
von: Daliri, Majid, et al.
Veröffentlicht: (2024)
von: Daliri, Majid, et al.
Veröffentlicht: (2024)
A Theoretical Perspective for Speculative Decoding Algorithm
von: Yin, Ming, et al.
Veröffentlicht: (2024)
von: Yin, Ming, et al.
Veröffentlicht: (2024)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
von: Hou, Yunlong, et al.
Veröffentlicht: (2025)
von: Hou, Yunlong, et al.
Veröffentlicht: (2025)
When Drafts Evolve: Speculative Decoding Meets Online Learning
von: Qian, Yu-Yang, et al.
Veröffentlicht: (2026)
von: Qian, Yu-Yang, et al.
Veröffentlicht: (2026)
Faster Cascades via Speculative Decoding
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
TS-DP: Reinforcement Speculative Decoding For Temporal Adaptive Diffusion Policy Acceleration
von: Li, Ye, et al.
Veröffentlicht: (2025)
von: Li, Ye, et al.
Veröffentlicht: (2025)
Self-Speculative Biased Decoding for Faster Re-Translation
von: Zeng, Linxiao, et al.
Veröffentlicht: (2025)
von: Zeng, Linxiao, et al.
Veröffentlicht: (2025)
AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
von: Jiang, Yuhua, et al.
Veröffentlicht: (2025)
von: Jiang, Yuhua, et al.
Veröffentlicht: (2025)
Scaling Capability in Token Space: An Analysis of Large Vision Language Model
von: Li, Tenghui, et al.
Veröffentlicht: (2024)
von: Li, Tenghui, et al.
Veröffentlicht: (2024)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
HiFloat4 Format for Language Model Inference
von: Luo, Yuanyong, et al.
Veröffentlicht: (2026)
von: Luo, Yuanyong, et al.
Veröffentlicht: (2026)
BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding
von: He, Liang, et al.
Veröffentlicht: (2026)
von: He, Liang, et al.
Veröffentlicht: (2026)
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
von: Wang, Yubo, et al.
Veröffentlicht: (2024)
von: Wang, Yubo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HiSpec: Hierarchical Speculative Decoding for LLMs
von: Kumar, Avinash, et al.
Veröffentlicht: (2025) -
CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative Drafter
von: Weng, Yepeng, et al.
Veröffentlicht: (2025) -
Recurrent Drafter for Fast Speculative Decoding in Large Language Models
von: Cheng, Yunfei, et al.
Veröffentlicht: (2024) -
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
von: Wang, Songsheng, et al.
Veröffentlicht: (2025) -
Mamba Drafters for Speculative Decoding
von: Choi, Daewon, et al.
Veröffentlicht: (2025)