SAGE: Accelerating Vision-Language Models via Entropy-Guided Adaptive Speculative Decoding

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tong, Yujia, Zhang, Tian, Wan, Yunyang, Lin, Kaiwei, Yuan, Jingling, Hu, Chuang
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911412976091136
author Tong, Yujia
Zhang, Tian
Wan, Yunyang
Lin, Kaiwei
Yuan, Jingling
Hu, Chuang
author_facet Tong, Yujia
Zhang, Tian
Wan, Yunyang
Lin, Kaiwei
Yuan, Jingling
Hu, Chuang
contents Speculative decoding has emerged as a promising approach to accelerate inference in vision-language models (VLMs) by enabling parallel verification of multiple draft tokens. However, existing methods rely on static tree structures that remain fixed throughout the decoding process, failing to adapt to the varying prediction difficulty across generation steps. This leads to suboptimal acceptance lengths and limited speedup. In this paper, we propose SAGE, a novel framework that dynamically adjusts the speculation tree structure based on real-time prediction uncertainty. Our key insight is that output entropy serves as a natural confidence indicator with strong temporal correlation across decoding steps. SAGE constructs deeper-narrower trees for high-confidence predictions to maximize speculation depth, and shallower-wider trees for uncertain predictions to diversify exploration. SAGE improves acceptance lengths and achieves faster acceleration compared to static tree baselines. Experiments on multiple benchmarks demonstrate the effectiveness of SAGE: without any loss in output quality, it delivers up to $3.36\times$ decoding speedup for LLaVA-OneVision-72B and $3.18\times$ for Qwen2.5-VL-72B.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00523
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SAGE: Accelerating Vision-Language Models via Entropy-Guided Adaptive Speculative Decoding
Tong, Yujia
Zhang, Tian
Wan, Yunyang
Lin, Kaiwei
Yuan, Jingling
Hu, Chuang
Computer Vision and Pattern Recognition
Speculative decoding has emerged as a promising approach to accelerate inference in vision-language models (VLMs) by enabling parallel verification of multiple draft tokens. However, existing methods rely on static tree structures that remain fixed throughout the decoding process, failing to adapt to the varying prediction difficulty across generation steps. This leads to suboptimal acceptance lengths and limited speedup. In this paper, we propose SAGE, a novel framework that dynamically adjusts the speculation tree structure based on real-time prediction uncertainty. Our key insight is that output entropy serves as a natural confidence indicator with strong temporal correlation across decoding steps. SAGE constructs deeper-narrower trees for high-confidence predictions to maximize speculation depth, and shallower-wider trees for uncertain predictions to diversify exploration. SAGE improves acceptance lengths and achieves faster acceleration compared to static tree baselines. Experiments on multiple benchmarks demonstrate the effectiveness of SAGE: without any loss in output quality, it delivers up to $3.36\times$ decoding speedup for LLaVA-OneVision-72B and $3.18\times$ for Qwen2.5-VL-72B.
title SAGE: Accelerating Vision-Language Models via Entropy-Guided Adaptive Speculative Decoding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.00523