Saved in:
Bibliographic Details
Main Authors: Wang, Zihua, Li, Ruibo, Du, Haozhe, Zhou, Joey Tianyi, Zhang, Yu, Yang, Xu
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.12728
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915769114165248
author Wang, Zihua
Li, Ruibo
Du, Haozhe
Zhou, Joey Tianyi
Zhang, Yu
Yang, Xu
author_facet Wang, Zihua
Li, Ruibo
Du, Haozhe
Zhou, Joey Tianyi
Zhang, Yu
Yang, Xu
contents Large language models and large multimodal models (LLMs and LMMs) deliver strong generative performance but suffer from slow decoding, a problem that becomes more severe when handling visual inputs, whose sequences typically contain many more tokens with lower information density than text. Speculative decoding accelerates LLM inference by letting a compact draft model propose candidate tokens that are selectively accepted by a larger target model, achieving speed-up without degrading quality. However, existing multimodal speculative decoding approaches largely ignore the structural characteristics of visual representations and usually rely on text-only draft models. In this paper, we introduce SpecFLASH, a speculative decoding framework tailored to LMMs that explicitly exploits multimodal structure when designing the draft model. We first mitigate redundancy in visual token sequences with a lightweight, latent-guided token compression module that compacts visual features while preserving semantics, and then leverage the co-occurrence and local correlations of visual entities via a semi-autoregressive decoding scheme that predicts multiple tokens in a single forward pass. Extensive experiments demonstrate that SpecFLASH consistently surpasses prior speculative decoding baselines, achieving up to $2.68\times$ speed-up on video captioning and $2.55\times$ on visual instruction tuning, relative to the original LMM. Our code is available here: https://github.com/ZihuaEvan/FlashSD/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12728
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SpecFLASH: A Latent-Guided Semi-autoregressive Speculative Decoding Framework for Efficient Multimodal Generation
Wang, Zihua
Li, Ruibo
Du, Haozhe
Zhou, Joey Tianyi
Zhang, Yu
Yang, Xu
Computer Vision and Pattern Recognition
Multimedia
Large language models and large multimodal models (LLMs and LMMs) deliver strong generative performance but suffer from slow decoding, a problem that becomes more severe when handling visual inputs, whose sequences typically contain many more tokens with lower information density than text. Speculative decoding accelerates LLM inference by letting a compact draft model propose candidate tokens that are selectively accepted by a larger target model, achieving speed-up without degrading quality. However, existing multimodal speculative decoding approaches largely ignore the structural characteristics of visual representations and usually rely on text-only draft models. In this paper, we introduce SpecFLASH, a speculative decoding framework tailored to LMMs that explicitly exploits multimodal structure when designing the draft model. We first mitigate redundancy in visual token sequences with a lightweight, latent-guided token compression module that compacts visual features while preserving semantics, and then leverage the co-occurrence and local correlations of visual entities via a semi-autoregressive decoding scheme that predicts multiple tokens in a single forward pass. Extensive experiments demonstrate that SpecFLASH consistently surpasses prior speculative decoding baselines, achieving up to $2.68\times$ speed-up on video captioning and $2.55\times$ on visual instruction tuning, relative to the original LMM. Our code is available here: https://github.com/ZihuaEvan/FlashSD/.
title SpecFLASH: A Latent-Guided Semi-autoregressive Speculative Decoding Framework for Efficient Multimodal Generation
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2505.12728