MMSpec: Benchmarking Speculative Decoding for Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shen, Hui, Wang, Xin, Zhang, Ping, Hsieh, Yunta, Han, Qi, Wan, Zhongwei, Zhang, Ziheng, Zhang, Jingxuan, Xiong, Jing, Liu, Ziyuan, Zhang, Yifan, Cao, Hangrui, Zhao, Chenyang, Zhang, Mi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912967861207040
author Shen, Hui
Wang, Xin
Zhang, Ping
Hsieh, Yunta
Han, Qi
Wan, Zhongwei
Zhang, Ziheng
Zhang, Jingxuan
Xiong, Jing
Liu, Ziyuan
Zhang, Yifan
Cao, Hangrui
Zhao, Chenyang
Zhang, Mi
author_facet Shen, Hui
Wang, Xin
Zhang, Ping
Hsieh, Yunta
Han, Qi
Wan, Zhongwei
Zhang, Ziheng
Zhang, Jingxuan
Xiong, Jing
Liu, Ziyuan
Zhang, Yifan
Cao, Hangrui
Zhao, Chenyang
Zhang, Mi
contents Vision-language models (VLMs) achieve strong performance on multimodal tasks but suffer from high inference latency due to large model sizes and long multimodal contexts. Speculative decoding has recently emerged as an effective acceleration technique, yet its behavior in VLMs remains insufficiently understood. We introduce MMSpec, the first benchmark for evaluating speculative decoding in vision-language models. MMSpec contains 600 multimodal samples across six task categories and integrates ten representative speculative decoding algorithms under a unified evaluation framework. Our study reveals three key findings: (1) methods designed for text-only LLMs degrade in multimodal scenarios, (2) vision awareness becomes increasingly important at larger batch sizes, and (3) throughput speedup alone does not reliably reflect latency performance. Motivated by these findings, we propose ViSkip, a plug-and-play speculative decoding method that dynamically adapts speculation to vision tokens and achieves state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2603_14989
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MMSpec: Benchmarking Speculative Decoding for Vision-Language Models
Shen, Hui
Wang, Xin
Zhang, Ping
Hsieh, Yunta
Han, Qi
Wan, Zhongwei
Zhang, Ziheng
Zhang, Jingxuan
Xiong, Jing
Liu, Ziyuan
Zhang, Yifan
Cao, Hangrui
Zhao, Chenyang
Zhang, Mi
Computer Vision and Pattern Recognition
Vision-language models (VLMs) achieve strong performance on multimodal tasks but suffer from high inference latency due to large model sizes and long multimodal contexts. Speculative decoding has recently emerged as an effective acceleration technique, yet its behavior in VLMs remains insufficiently understood. We introduce MMSpec, the first benchmark for evaluating speculative decoding in vision-language models. MMSpec contains 600 multimodal samples across six task categories and integrates ten representative speculative decoding algorithms under a unified evaluation framework. Our study reveals three key findings: (1) methods designed for text-only LLMs degrade in multimodal scenarios, (2) vision awareness becomes increasingly important at larger batch sizes, and (3) throughput speedup alone does not reliably reflect latency performance. Motivated by these findings, we propose ViSkip, a plug-and-play speculative decoding method that dynamically adapts speculation to vision tokens and achieves state-of-the-art performance.
title MMSpec: Benchmarking Speculative Decoding for Vision-Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.14989