CoVSpec: Efficient Device-Edge Co-Inference for Vision-Language Models via Speculative Decoding
Fuente:
arXiv
Salvato in:
| Autori principali: | Jia, Yuanyuan, Tang, Shunpu, Yang, Qianqian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FeDeRA:Efficient Fine-tuning of Language Models in Federated Learning Leveraging Weight Decomposition
di: Yan, Yuxuan, et al.
Pubblicazione: (2024)
di: Yan, Yuxuan, et al.
Pubblicazione: (2024)
GCoDE: Efficient Device-Edge Co-Inference for GNNs via Architecture-Mapping Co-Search
di: Zhou, Ao, et al.
Pubblicazione: (2025)
di: Zhou, Ao, et al.
Pubblicazione: (2025)
WISV: Wireless-Informed Semantic Verification for Distributed Speculative Decoding in Device-Edge LLM Inference
di: Liu, Zixuan, et al.
Pubblicazione: (2026)
di: Liu, Zixuan, et al.
Pubblicazione: (2026)
FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
di: Bajpai, Divya Jyoti, et al.
Pubblicazione: (2025)
di: Bajpai, Divya Jyoti, et al.
Pubblicazione: (2025)
Dynamic-Width Speculative Beam Decoding for Efficient LLM Inference
di: Qin, Zongyue, et al.
Pubblicazione: (2024)
di: Qin, Zongyue, et al.
Pubblicazione: (2024)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
di: Huang, Haiduo, et al.
Pubblicazione: (2025)
di: Huang, Haiduo, et al.
Pubblicazione: (2025)
Graph Neural Networks Automated Design and Deployment on Device-Edge Co-Inference Systems
di: Zhou, Ao, et al.
Pubblicazione: (2024)
di: Zhou, Ao, et al.
Pubblicazione: (2024)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
di: Liu, Xing, et al.
Pubblicazione: (2025)
di: Liu, Xing, et al.
Pubblicazione: (2025)
Conformal Sparsification for Bandwidth-Efficient Edge-Cloud Speculative Decoding
di: Bhattacharjee, Payel, et al.
Pubblicazione: (2025)
di: Bhattacharjee, Payel, et al.
Pubblicazione: (2025)
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
di: Park, Jihoon, et al.
Pubblicazione: (2025)
di: Park, Jihoon, et al.
Pubblicazione: (2025)
HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
di: Yang, Ze, et al.
Pubblicazione: (2024)
di: Yang, Ze, et al.
Pubblicazione: (2024)
Speculative Decoding for Multi-Sample Inference
di: Li, Yiwei, et al.
Pubblicazione: (2025)
di: Li, Yiwei, et al.
Pubblicazione: (2025)
Efficient Partitioning Vision Transformer on Edge Devices for Distributed Inference
di: Liu, Xiang, et al.
Pubblicazione: (2024)
di: Liu, Xiang, et al.
Pubblicazione: (2024)
Fast Large Language Model Collaborative Decoding via Speculation
di: Fu, Jiale, et al.
Pubblicazione: (2025)
di: Fu, Jiale, et al.
Pubblicazione: (2025)
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
di: Wang, Songsheng, et al.
Pubblicazione: (2025)
di: Wang, Songsheng, et al.
Pubblicazione: (2025)
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
di: Li, Guanghao, et al.
Pubblicazione: (2025)
di: Li, Guanghao, et al.
Pubblicazione: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
di: Wen, Zhuofan, et al.
Pubblicazione: (2024)
di: Wen, Zhuofan, et al.
Pubblicazione: (2024)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
di: Zhao, Yilong, et al.
Pubblicazione: (2025)
di: Zhao, Yilong, et al.
Pubblicazione: (2025)
HIPPO: Accelerating Video Large Language Models Inference via Holistic-aware Parallel Speculative Decoding
di: Lv, Qitan, et al.
Pubblicazione: (2026)
di: Lv, Qitan, et al.
Pubblicazione: (2026)
Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding
di: Ryu, Hyun, et al.
Pubblicazione: (2024)
di: Ryu, Hyun, et al.
Pubblicazione: (2024)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
di: Jeon, Wonseok, et al.
Pubblicazione: (2024)
di: Jeon, Wonseok, et al.
Pubblicazione: (2024)
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding
di: Bang, Jehyeon, et al.
Pubblicazione: (2026)
di: Bang, Jehyeon, et al.
Pubblicazione: (2026)
HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
di: Xie, Zhinan, et al.
Pubblicazione: (2025)
di: Xie, Zhinan, et al.
Pubblicazione: (2025)
On Speculative Decoding for Multimodal Large Language Models
di: Gagrani, Mukul, et al.
Pubblicazione: (2024)
di: Gagrani, Mukul, et al.
Pubblicazione: (2024)
S$^4$C: Speculative Sampling with Syntactic and Semantic Coherence for Efficient Inference of Large Language Models
di: He, Tao, et al.
Pubblicazione: (2025)
di: He, Tao, et al.
Pubblicazione: (2025)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
di: Lin, Zijian, et al.
Pubblicazione: (2025)
di: Lin, Zijian, et al.
Pubblicazione: (2025)
Efficient Long CoT Reasoning in Small Language Models
di: Wang, Zhaoyang, et al.
Pubblicazione: (2025)
di: Wang, Zhaoyang, et al.
Pubblicazione: (2025)
Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding
di: Wang, Ziyao, et al.
Pubblicazione: (2025)
di: Wang, Ziyao, et al.
Pubblicazione: (2025)
Confidence-Modulated Speculative Decoding for Large Language Models
di: Sen, Jaydip, et al.
Pubblicazione: (2025)
di: Sen, Jaydip, et al.
Pubblicazione: (2025)
Speculative Decoding Reimagined for Multimodal Large Language Models
di: Lin, Luxi, et al.
Pubblicazione: (2025)
di: Lin, Luxi, et al.
Pubblicazione: (2025)
Collaborative Edge-to-Server Inference for Vision-Language Models
di: Song, Soochang, et al.
Pubblicazione: (2025)
di: Song, Soochang, et al.
Pubblicazione: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
di: Liao, Baohao, et al.
Pubblicazione: (2025)
di: Liao, Baohao, et al.
Pubblicazione: (2025)
CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models
di: Cao, Zongsheng, et al.
Pubblicazione: (2025)
di: Cao, Zongsheng, et al.
Pubblicazione: (2025)
Encoder-Decoder Diffusion Language Models for Efficient Training and Inference
di: Arriola, Marianne, et al.
Pubblicazione: (2025)
di: Arriola, Marianne, et al.
Pubblicazione: (2025)
Plato: Plan to Efficiently Decode for Large Language Model Inference
di: Jin, Shuowei, et al.
Pubblicazione: (2024)
di: Jin, Shuowei, et al.
Pubblicazione: (2024)
CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
Large Language Models as Co-Pilots for Causal Inference in Medical Studies
di: Alaa, Ahmed, et al.
Pubblicazione: (2024)
di: Alaa, Ahmed, et al.
Pubblicazione: (2024)
Clover-2: Accurate Inference for Regressive Lightweight Speculative Decoding
di: Xiao, Bin, et al.
Pubblicazione: (2024)
di: Xiao, Bin, et al.
Pubblicazione: (2024)
Faster Cascades via Speculative Decoding
di: Narasimhan, Harikrishna, et al.
Pubblicazione: (2024)
di: Narasimhan, Harikrishna, et al.
Pubblicazione: (2024)
Efficient Adaptive Rejection Sampling for Accelerating Speculative Decoding in Large Language Models
di: Sun, Chendong, et al.
Pubblicazione: (2025)
di: Sun, Chendong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FeDeRA:Efficient Fine-tuning of Language Models in Federated Learning Leveraging Weight Decomposition
di: Yan, Yuxuan, et al.
Pubblicazione: (2024) -
GCoDE: Efficient Device-Edge Co-Inference for GNNs via Architecture-Mapping Co-Search
di: Zhou, Ao, et al.
Pubblicazione: (2025) -
WISV: Wireless-Informed Semantic Verification for Distributed Speculative Decoding in Device-Edge LLM Inference
di: Liu, Zixuan, et al.
Pubblicazione: (2026) -
FastVLM: Self-Speculative Decoding for Fast Vision-Language Model Inference
di: Bajpai, Divya Jyoti, et al.
Pubblicazione: (2025) -
Dynamic-Width Speculative Beam Decoding for Efficient LLM Inference
di: Qin, Zongyue, et al.
Pubblicazione: (2024)