SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Svirschevski, Ruslan, May, Avner, Chen, Zhuoming, Chen, Beidi, Jia, Zhihao, Ryabinin, Max |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
von: Chen, Zhuoming, et al.
Veröffentlicht: (2024)
von: Chen, Zhuoming, et al.
Veröffentlicht: (2024)
AutoJudge: Judge Decoding Without Manual Annotation
von: Garipov, Roman, et al.
Veröffentlicht: (2025)
von: Garipov, Roman, et al.
Veröffentlicht: (2025)
MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
von: Sadhukhan, Ranajoy, et al.
Veröffentlicht: (2024)
von: Sadhukhan, Ranajoy, et al.
Veröffentlicht: (2024)
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
Nearest Neighbor Speculative Decoding for LLM Generation and Attribution
von: Li, Minghan, et al.
Veröffentlicht: (2024)
von: Li, Minghan, et al.
Veröffentlicht: (2024)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025)
von: Li, Zikun, et al.
Veröffentlicht: (2025)
SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting
von: Shi, Weijie, et al.
Veröffentlicht: (2026)
von: Shi, Weijie, et al.
Veröffentlicht: (2026)
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
von: Sun, Ryan, et al.
Veröffentlicht: (2024)
von: Sun, Ryan, et al.
Veröffentlicht: (2024)
NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context Vocabularies
von: Chen, Zhiyang, et al.
Veröffentlicht: (2026)
von: Chen, Zhiyang, et al.
Veröffentlicht: (2026)
HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative Decoding
von: Liu, Siran, et al.
Veröffentlicht: (2025)
von: Liu, Siran, et al.
Veröffentlicht: (2025)
HiSpec: Hierarchical Speculative Decoding for LLMs
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
Spec-LLaVA: Accelerating Vision-Language Models with Dynamic Tree-Based Speculative Decoding
von: Huo, Mingxiao, et al.
Veröffentlicht: (2025)
von: Huo, Mingxiao, et al.
Veröffentlicht: (2025)
ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
von: Kang, Jialiang, et al.
Veröffentlicht: (2025)
von: Kang, Jialiang, et al.
Veröffentlicht: (2025)
DART: Diffusion-Inspired Speculative Decoding for Fast LLM Inference
von: Liu, Fuliang, et al.
Veröffentlicht: (2026)
von: Liu, Fuliang, et al.
Veröffentlicht: (2026)
Sirius: Contextual Sparsity with Correction for Efficient LLMs
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
LogitSpec: Accelerating Retrieval-based Speculative Decoding via Next Next Token Speculation
von: Liu, Tianyu, et al.
Veröffentlicht: (2025)
von: Liu, Tianyu, et al.
Veröffentlicht: (2025)
Tutorial Proposal: Speculative Decoding for Efficient LLM Inference
von: Xia, Heming, et al.
Veröffentlicht: (2025)
von: Xia, Heming, et al.
Veröffentlicht: (2025)
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
MagicPIG: LSH Sampling for Efficient LLM Generation
von: Chen, Zhuoming, et al.
Veröffentlicht: (2024)
von: Chen, Zhuoming, et al.
Veröffentlicht: (2024)
GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
SpecTr-GBV: Multi-Draft Block Verification Accelerating Speculative Decoding
von: Lin, Yijun, et al.
Veröffentlicht: (2026)
von: Lin, Yijun, et al.
Veröffentlicht: (2026)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2024)
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2024)
EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation
von: Zhang, Shuyu, et al.
Veröffentlicht: (2026)
von: Zhang, Shuyu, et al.
Veröffentlicht: (2026)
Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration
von: Xia, Heming, et al.
Veröffentlicht: (2024)
von: Xia, Heming, et al.
Veröffentlicht: (2024)
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
Speculative Speculative Decoding
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
PSD: Pushing the Pareto Frontier of Diffusion LLMs via Parallel Speculative Decoding
von: Sun, Shengyin, et al.
Veröffentlicht: (2026)
von: Sun, Shengyin, et al.
Veröffentlicht: (2026)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding
von: Li, Shenggui, et al.
Veröffentlicht: (2026)
von: Li, Shenggui, et al.
Veröffentlicht: (2026)
Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters
von: Yi, Euiin, et al.
Veröffentlicht: (2024)
von: Yi, Euiin, et al.
Veröffentlicht: (2024)
SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
von: Chen, Zhuoming, et al.
Veröffentlicht: (2024) -
AutoJudge: Judge Decoding Without Manual Annotation
von: Garipov, Roman, et al.
Veröffentlicht: (2025) -
MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
von: Sadhukhan, Ranajoy, et al.
Veröffentlicht: (2024) -
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
von: Sun, Hanshi, et al.
Veröffentlicht: (2024) -
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)