Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wang, Zhibin, Zhong, Ziyu, Shen, Nuo, Zhou, Yuhang, Gu, Rong, Zhong, Sheng
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2605.19893
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911702088417280
author Wang, Zhibin
Zhong, Ziyu
Shen, Nuo
Zhou, Yuhang
Gu, Rong
Zhong, Sheng
author_facet Wang, Zhibin
Zhong, Ziyu
Shen, Nuo
Zhou, Yuhang
Gu, Rong
Zhong, Sheng
contents Speculative decoding and dynamic sparse attention are two complementary approaches for accelerating long-context LLM inference: the former amortizes target-model execution across multiple verifier queries, while the latter reduces each query's KV-cache working set. Directly combining them, however, exposes a structural mismatch: speculative verification relies on cross-query commonality, whereas dynamic sparse attention assigns query-specific sparse layouts. This mismatch limits KV-block reuse, amplifies NSA's branch-wise overheads, and makes verification strategy selection input- and regime-dependent. We present SSV, a sparse speculative-verification framework that turns dynamic sparse attention into a verification-oriented workload. SSV combines overlap-aware grouped-query execution, refresh/reuse-based NSA kernel fusion, and profile-guided prompt-adaptive orchestration to improve cross-query reuse, reduce selected-index and branch-fusion overheads, and select effective draft-verification strategies under user-specified precision classes. Experiments on NVIDIA H100 GPUs show that SSV achieves up to 3.49x end-to-end throughput over autoregressive NSA decoding and up to 6.86x kernel speedups for sparse speculative verification.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19893
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SSV: Sparse Speculative Verification for Efficient LLM Inference
Wang, Zhibin
Zhong, Ziyu
Shen, Nuo
Zhou, Yuhang
Gu, Rong
Zhong, Sheng
Operating Systems
Speculative decoding and dynamic sparse attention are two complementary approaches for accelerating long-context LLM inference: the former amortizes target-model execution across multiple verifier queries, while the latter reduces each query's KV-cache working set. Directly combining them, however, exposes a structural mismatch: speculative verification relies on cross-query commonality, whereas dynamic sparse attention assigns query-specific sparse layouts. This mismatch limits KV-block reuse, amplifies NSA's branch-wise overheads, and makes verification strategy selection input- and regime-dependent. We present SSV, a sparse speculative-verification framework that turns dynamic sparse attention into a verification-oriented workload. SSV combines overlap-aware grouped-query execution, refresh/reuse-based NSA kernel fusion, and profile-guided prompt-adaptive orchestration to improve cross-query reuse, reduce selected-index and branch-fusion overheads, and select effective draft-verification strategies under user-specified precision classes. Experiments on NVIDIA H100 GPUs show that SSV achieves up to 3.49x end-to-end throughput over autoregressive NSA decoding and up to 6.86x kernel speedups for sparse speculative verification.
title SSV: Sparse Speculative Verification for Efficient LLM Inference
topic Operating Systems
url https://arxiv.org/abs/2605.19893