Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Zhuoming, May, Avner, Svirschevski, Ruslan, Huang, Yuhsun, Ryabinin, Max, Jia, Zhihao, Chen, Beidi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
von: Svirschevski, Ruslan, et al.
Veröffentlicht: (2024)
von: Svirschevski, Ruslan, et al.
Veröffentlicht: (2024)
AutoJudge: Judge Decoding Without Manual Annotation
von: Garipov, Roman, et al.
Veröffentlicht: (2025)
von: Garipov, Roman, et al.
Veröffentlicht: (2025)
MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
von: Sadhukhan, Ranajoy, et al.
Veröffentlicht: (2024)
von: Sadhukhan, Ranajoy, et al.
Veröffentlicht: (2024)
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting
von: Lv, Kai, et al.
Veröffentlicht: (2025)
von: Lv, Kai, et al.
Veröffentlicht: (2025)
Nearest Neighbor Speculative Decoding for LLM Generation and Attribution
von: Li, Minghan, et al.
Veröffentlicht: (2024)
von: Li, Minghan, et al.
Veröffentlicht: (2024)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025)
von: Li, Zikun, et al.
Veröffentlicht: (2025)
Sirius: Contextual Sparsity with Correction for Efficient LLMs
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
von: Zhou, Yang, et al.
Veröffentlicht: (2024)
GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
von: Zhou, Yang, et al.
Veröffentlicht: (2025)
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
von: Liu, Jingyu, et al.
Veröffentlicht: (2025)
Speculative Speculative Decoding
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
von: Oliaro, Gabriele, et al.
Veröffentlicht: (2024)
Multi-Candidate Speculative Decoding
von: Yang, Sen, et al.
Veröffentlicht: (2024)
von: Yang, Sen, et al.
Veröffentlicht: (2024)
Kinetics: Rethinking Test-Time Scaling Laws
von: Sadhukhan, Ranajoy, et al.
Veröffentlicht: (2025)
von: Sadhukhan, Ranajoy, et al.
Veröffentlicht: (2025)
Mind Your Format: Towards Consistent Evaluation of In-Context Learning Improvements
von: Voronov, Anton, et al.
Veröffentlicht: (2024)
von: Voronov, Anton, et al.
Veröffentlicht: (2024)
Speculative Contrastive Decoding
von: Yuan, Hongyi, et al.
Veröffentlicht: (2023)
von: Yuan, Hongyi, et al.
Veröffentlicht: (2023)
MagicPIG: LSH Sampling for Efficient LLM Generation
von: Chen, Zhuoming, et al.
Veröffentlicht: (2024)
von: Chen, Zhuoming, et al.
Veröffentlicht: (2024)
Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative Decoding
von: Wang, Pei-Shuo, et al.
Veröffentlicht: (2025)
von: Wang, Pei-Shuo, et al.
Veröffentlicht: (2025)
Speculative Decoding with a Speculative Vocabulary
von: Williams, Miles, et al.
Veröffentlicht: (2026)
von: Williams, Miles, et al.
Veröffentlicht: (2026)
SSSD: Simply-Scalable Speculative Decoding
von: Marzollo, Michele, et al.
Veröffentlicht: (2024)
von: Marzollo, Michele, et al.
Veröffentlicht: (2024)
DFlash: Block Diffusion for Flash Speculative Decoding
von: Chen, Jian, et al.
Veröffentlicht: (2026)
von: Chen, Jian, et al.
Veröffentlicht: (2026)
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2024)
von: Elhoushi, Mostafa, et al.
Veröffentlicht: (2024)
Label Privacy in Split Learning for Large Models with Parameter-Efficient Training
von: Zmushko, Philip, et al.
Veröffentlicht: (2024)
von: Zmushko, Philip, et al.
Veröffentlicht: (2024)
Decoding Speculative Decoding
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
A Theoretical Perspective for Speculative Decoding Algorithm
von: Yin, Ming, et al.
Veröffentlicht: (2024)
von: Yin, Ming, et al.
Veröffentlicht: (2024)
HADES: Hardware Accelerated Decoding for Efficient Speculation in Large Language Models
von: Yang, Ze, et al.
Veröffentlicht: (2024)
von: Yang, Ze, et al.
Veröffentlicht: (2024)
SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting
von: Shi, Weijie, et al.
Veröffentlicht: (2026)
von: Shi, Weijie, et al.
Veröffentlicht: (2026)
TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
von: Yang, Lijie, et al.
Veröffentlicht: (2024)
von: Yang, Lijie, et al.
Veröffentlicht: (2024)
Scalable LLM Reasoning Acceleration with Low-rank Distillation
von: Dong, Harry, et al.
Veröffentlicht: (2025)
von: Dong, Harry, et al.
Veröffentlicht: (2025)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
Graph-Structured Speculative Decoding
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2024)
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2024)
3-Model Speculative Decoding
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
Towards Optimal Multi-draft Speculative Decoding
von: Hu, Zhengmian, et al.
Veröffentlicht: (2025)
von: Hu, Zhengmian, et al.
Veröffentlicht: (2025)
Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding
von: Kim, Sungkyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungkyun, et al.
Veröffentlicht: (2025)
Performance-Driven Policy Optimization for Speculative Decoding with Adaptive Windowing
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
von: Jiang, Jie, et al.
Veröffentlicht: (2026)
RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding
von: Chen, Guanzheng, et al.
Veröffentlicht: (2025)
von: Chen, Guanzheng, et al.
Veröffentlicht: (2025)
Dynamic Speculation Lookahead Accelerates Speculative Decoding of Large Language Models
von: Mamou, Jonathan, et al.
Veröffentlicht: (2024)
von: Mamou, Jonathan, et al.
Veröffentlicht: (2024)
PSD: Pushing the Pareto Frontier of Diffusion LLMs via Parallel Speculative Decoding
von: Sun, Shengyin, et al.
Veröffentlicht: (2026)
von: Sun, Shengyin, et al.
Veröffentlicht: (2026)
Improving Multi-candidate Speculative Decoding
von: Lu, Xiaofan, et al.
Veröffentlicht: (2024)
von: Lu, Xiaofan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
von: Svirschevski, Ruslan, et al.
Veröffentlicht: (2024) -
AutoJudge: Judge Decoding Without Manual Annotation
von: Garipov, Roman, et al.
Veröffentlicht: (2025) -
MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding
von: Sadhukhan, Ranajoy, et al.
Veröffentlicht: (2024) -
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
von: Sun, Hanshi, et al.
Veröffentlicht: (2024) -
DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting
von: Lv, Kai, et al.
Veröffentlicht: (2025)