MMSpec: Benchmarking Speculative Decoding for Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Hui, Wang, Xin, Zhang, Ping, Hsieh, Yunta, Han, Qi, Wan, Zhongwei, Zhang, Ziheng, Zhang, Jingxuan, Xiong, Jing, Liu, Ziyuan, Zhang, Yifan, Cao, Hangrui, Zhao, Chenyang, Zhang, Mi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models
by: Zhang, Jingxuan, et al.
Published: (2026)
by: Zhang, Jingxuan, et al.
Published: (2026)
MathGen: Revealing the Illusion of Mathematical Competence through Text-to-Image Generation
by: Liu, Ruiyao, et al.
Published: (2026)
by: Liu, Ruiyao, et al.
Published: (2026)
Famba-V: Fast Vision Mamba with Cross-Layer Token Fusion
by: Shen, Hui, et al.
Published: (2024)
by: Shen, Hui, et al.
Published: (2024)
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills
by: Lei, Yingtie, et al.
Published: (2026)
by: Lei, Yingtie, et al.
Published: (2026)
MMFormalizer: Multimodal Autoformalization in the Wild
by: Xiong, Jing, et al.
Published: (2026)
by: Xiong, Jing, et al.
Published: (2026)
SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model Compression
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Argus: Benchmarking and Enhancing Vision-Language Models for 3D Radiology Report Generation
by: Liu, Che, et al.
Published: (2024)
by: Liu, Che, et al.
Published: (2024)
SAM Decoding: Speculative Decoding via Suffix Automaton
by: Hu, Yuxuan, et al.
Published: (2024)
by: Hu, Yuxuan, et al.
Published: (2024)
MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference
by: Wan, Zhongwei, et al.
Published: (2025)
by: Wan, Zhongwei, et al.
Published: (2025)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
by: Zhang, Zihong, et al.
Published: (2026)
by: Zhang, Zihong, et al.
Published: (2026)
SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
by: Zhang, Ziyi, et al.
Published: (2025)
by: Zhang, Ziyi, et al.
Published: (2025)
MEIT: Multimodal Electrocardiogram Instruction Tuning on Large Language Models for Report Generation
by: Wan, Zhongwei, et al.
Published: (2024)
by: Wan, Zhongwei, et al.
Published: (2024)
PACER: Blockwise Pre-verification for Speculative Decoding with Adaptive Length
by: Zhang, Situo, et al.
Published: (2026)
by: Zhang, Situo, et al.
Published: (2026)
SAGE: Accelerating Vision-Language Models via Entropy-Guided Adaptive Speculative Decoding
by: Tong, Yujia, et al.
Published: (2026)
by: Tong, Yujia, et al.
Published: (2026)
Efficient Diffusion Models: A Survey
by: Shen, Hui, et al.
Published: (2025)
by: Shen, Hui, et al.
Published: (2025)
ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
by: Hu, Xinyi, et al.
Published: (2026)
by: Hu, Xinyi, et al.
Published: (2026)
PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
by: Shen, Hui, et al.
Published: (2025)
by: Shen, Hui, et al.
Published: (2025)
The Internet of Things in the Era of Generative AI: Vision and Challenges
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Speculative Decoding for Autoregressive Video Generation
by: Hu, Yuezhou, et al.
Published: (2026)
by: Hu, Yuezhou, et al.
Published: (2026)
DSDR: Dual-Scale Diversity Regularization for Exploration in LLM Reasoning
by: Wan, Zhongwei, et al.
Published: (2026)
by: Wan, Zhongwei, et al.
Published: (2026)
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
by: Wan, Zhongwei, et al.
Published: (2025)
by: Wan, Zhongwei, et al.
Published: (2025)
Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation
by: Li, Xingyao, et al.
Published: (2026)
by: Li, Xingyao, et al.
Published: (2026)
Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
by: Li, Ruanjun, et al.
Published: (2025)
by: Li, Ruanjun, et al.
Published: (2025)
Batch Speculative Decoding Done Right
by: Zhang, Ranran Haoran, et al.
Published: (2025)
by: Zhang, Ranran Haoran, et al.
Published: (2025)
Continuous Speculative Decoding for Autoregressive Image Generation
by: Wang, Zili, et al.
Published: (2024)
by: Wang, Zili, et al.
Published: (2024)
When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding?
by: Liu, Tianyu, et al.
Published: (2026)
by: Liu, Tianyu, et al.
Published: (2026)
Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning
by: Zhang, Jiebin, et al.
Published: (2026)
by: Zhang, Jiebin, et al.
Published: (2026)
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
by: Shen, Yuhao, et al.
Published: (2026)
by: Shen, Yuhao, et al.
Published: (2026)
DySpec: Faster Speculative Decoding with Dynamic Token Tree Structure
by: Xiong, Yunfan, et al.
Published: (2024)
by: Xiong, Yunfan, et al.
Published: (2024)
Enhancing Test-Time Scaling of Large Language Models with Hierarchical Retrieval-Augmented MCTS
by: Dou, Alex ZH, et al.
Published: (2025)
by: Dou, Alex ZH, et al.
Published: (2025)
Online Speculative Decoding
by: Liu, Xiaoxuan, et al.
Published: (2023)
by: Liu, Xiaoxuan, et al.
Published: (2023)
Scaling Laws for Speculative Decoding
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
by: Huang, Jianuo, et al.
Published: (2026)
by: Huang, Jianuo, et al.
Published: (2026)
KOALA: Enhancing Speculative Decoding for LLM via Multi-Layer Draft Heads with Adversarial Learning
by: Zhang, Kaiqi, et al.
Published: (2024)
by: Zhang, Kaiqi, et al.
Published: (2024)
Reinforcement Speculative Decoding for Fast Ranking
by: Du, Yingpeng, et al.
Published: (2025)
by: Du, Yingpeng, et al.
Published: (2025)
Self Speculative Decoding for Diffusion Large Language Models
by: Gao, Yifeng, et al.
Published: (2025)
by: Gao, Yifeng, et al.
Published: (2025)
Recurrent Drafter for Fast Speculative Decoding in Large Language Models
by: Cheng, Yunfei, et al.
Published: (2024)
by: Cheng, Yunfei, et al.
Published: (2024)
Fast Collaborative Inference via Distributed Speculative Decoding
by: Zheng, Ce, et al.
Published: (2025)
by: Zheng, Ce, et al.
Published: (2025)
VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs
by: Goel, Raghavv, et al.
Published: (2025)
by: Goel, Raghavv, et al.
Published: (2025)
Similar Items
-
QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models
by: Zhang, Jingxuan, et al.
Published: (2026) -
MathGen: Revealing the Illusion of Mathematical Competence through Text-to-Image Generation
by: Liu, Ruiyao, et al.
Published: (2026) -
Famba-V: Fast Vision Mamba with Cross-Layer Token Fusion
by: Shen, Hui, et al.
Published: (2024) -
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills
by: Lei, Yingtie, et al.
Published: (2026) -
MMFormalizer: Multimodal Autoformalization in the Wild
by: Xiong, Jing, et al.
Published: (2026)