Speculative Decoding: Performance or Illusion?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xiaoxuan, Yu, Jiaxiang, Park, Jongseok, Stoica, Ion, Cheung, Alvin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
Qrita: High-performance Top-k and Top-p using Pivot-based Truncation and Selection
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
von: Cemri, Mert, et al.
Veröffentlicht: (2025)
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2024)
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding
von: Su, Xin, et al.
Veröffentlicht: (2026)
von: Su, Xin, et al.
Veröffentlicht: (2026)
The Disparate Impacts of Speculative Decoding
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
Cross-Attention Speculative Decoding
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
Mamba Drafters for Speculative Decoding
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
Constrained Decoding with Speculative Lookaheads
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding
von: Jin, Tao, et al.
Veröffentlicht: (2026)
von: Jin, Tao, et al.
Veröffentlicht: (2026)
CREST: Effectively Compacting a Datastore For Retrieval-Based Speculative Decoding
von: Ho, Sophia, et al.
Veröffentlicht: (2024)
von: Ho, Sophia, et al.
Veröffentlicht: (2024)
Batch Speculative Decoding Done Right
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
RASD: Retrieval-Augmented Speculative Decoding
von: Quan, Guofeng, et al.
Veröffentlicht: (2025)
von: Quan, Guofeng, et al.
Veröffentlicht: (2025)
Cacheback: Speculative Decoding With Nothing But Cache
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
On Speculative Decoding for Multimodal Large Language Models
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
PACER: Blockwise Pre-verification for Speculative Decoding with Adaptive Length
von: Zhang, Situo, et al.
Veröffentlicht: (2026)
von: Zhang, Situo, et al.
Veröffentlicht: (2026)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
Traversal Verification for Speculative Tree Decoding
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
OR-Bench: An Over-Refusal Benchmark for Large Language Models
von: Cui, Justin, et al.
Veröffentlicht: (2024)
von: Cui, Justin, et al.
Veröffentlicht: (2024)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
Acceptance Dynamics Across Cognitive Domains in Speculative Decoding
von: Mahmoud, Saif
Veröffentlicht: (2026)
von: Mahmoud, Saif
Veröffentlicht: (2026)
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
von: An, Zihao, et al.
Veröffentlicht: (2026)
von: An, Zihao, et al.
Veröffentlicht: (2026)
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
When, What, and How: Rethinking Retrieval-Enhanced Speculative Decoding
von: Fang, Min, et al.
Veröffentlicht: (2025)
von: Fang, Min, et al.
Veröffentlicht: (2025)
Entropy-Aware Speculative Decoding Toward Improved LLM Reasoning
von: Su, Tiancheng, et al.
Veröffentlicht: (2025)
von: Su, Tiancheng, et al.
Veröffentlicht: (2025)
Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
ClawEnvKit: Automatic Environment Generation for Claw-Like Agents
von: Li, Xirui, et al.
Veröffentlicht: (2026)
von: Li, Xirui, et al.
Veröffentlicht: (2026)
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
von: Sun, Ryan, et al.
Veröffentlicht: (2024)
von: Sun, Ryan, et al.
Veröffentlicht: (2024)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
Faster Cascades via Speculative Decoding
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures
von: Zhang, Situo, et al.
Veröffentlicht: (2024)
von: Zhang, Situo, et al.
Veröffentlicht: (2024)
Reasons and Solutions for the Decline in Model Performance after Editing
von: Huang, Xiusheng, et al.
Veröffentlicht: (2024)
von: Huang, Xiusheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023) -
Qrita: High-performance Top-k and Top-p using Pivot-based Truncation and Selection
von: Park, Jongseok, et al.
Veröffentlicht: (2026) -
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
von: Park, Jongseok, et al.
Veröffentlicht: (2026) -
$\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Drafts
von: Cemri, Mert, et al.
Veröffentlicht: (2025) -
TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2024)