Tutorial Proposal: Speculative Decoding for Efficient LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xia, Heming, Du, Cunxiao, Li, Yongqi, Liu, Qian, Li, Wenjie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration
von: Xia, Heming, et al.
Veröffentlicht: (2024)
von: Xia, Heming, et al.
Veröffentlicht: (2024)
ToolSpec: Accelerating Tool Calling via Schema-Aware and Retrieval-Augmented Speculative Decoding
von: Xia, Heming, et al.
Veröffentlicht: (2026)
von: Xia, Heming, et al.
Veröffentlicht: (2026)
Merlin's Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting
von: Xia, Heming, et al.
Veröffentlicht: (2025)
von: Xia, Heming, et al.
Veröffentlicht: (2025)
Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
von: Xia, Heming, et al.
Veröffentlicht: (2024)
von: Xia, Heming, et al.
Veröffentlicht: (2024)
GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative Decoding
von: Du, Cunxiao, et al.
Veröffentlicht: (2024)
von: Du, Cunxiao, et al.
Veröffentlicht: (2024)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
Efficient Inference for Large Language Model-based Generative Recommendation
von: Lin, Xinyu, et al.
Veröffentlicht: (2024)
von: Lin, Xinyu, et al.
Veröffentlicht: (2024)
Enhancing Tool Retrieval with Iterative Feedback from Large Language Models
von: Xu, Qiancheng, et al.
Veröffentlicht: (2024)
von: Xu, Qiancheng, et al.
Veröffentlicht: (2024)
PEToolLLM: Towards Personalized Tool Learning in Large Language Models
von: Xu, Qiancheng, et al.
Veröffentlicht: (2025)
von: Xu, Qiancheng, et al.
Veröffentlicht: (2025)
KNN-SSD: Enabling Dynamic Self-Speculative Decoding via Nearest Neighbor Layer Set Optimization
von: Song, Mingbo, et al.
Veröffentlicht: (2025)
von: Song, Mingbo, et al.
Veröffentlicht: (2025)
TokenSkip: Controllable Chain-of-Thought Compression in LLMs
von: Xia, Heming, et al.
Veröffentlicht: (2025)
von: Xia, Heming, et al.
Veröffentlicht: (2025)
DART: Diffusion-Inspired Speculative Decoding for Fast LLM Inference
von: Liu, Fuliang, et al.
Veröffentlicht: (2026)
von: Liu, Fuliang, et al.
Veröffentlicht: (2026)
Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters
von: Yi, Euiin, et al.
Veröffentlicht: (2024)
von: Yi, Euiin, et al.
Veröffentlicht: (2024)
SDSAT: Accelerating LLM Inference through Speculative Decoding with Semantic Adaptive Tokens
von: Liu, Chengbo, et al.
Veröffentlicht: (2024)
von: Liu, Chengbo, et al.
Veröffentlicht: (2024)
Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference
von: Zhou, Xuwen, et al.
Veröffentlicht: (2026)
von: Zhou, Xuwen, et al.
Veröffentlicht: (2026)
Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
AdaSD: Adaptive Speculative Decoding for Efficient Language Model Inference
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation
von: Li, Xingyao, et al.
Veröffentlicht: (2026)
von: Li, Xingyao, et al.
Veröffentlicht: (2026)
SpecVLM: Enhancing Speculative Decoding of Video LLMs via Verifier-Guided Token Pruning
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
von: Ji, Yicheng, et al.
Veröffentlicht: (2025)
RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding
von: Chen, Guanzheng, et al.
Veröffentlicht: (2025)
von: Chen, Guanzheng, et al.
Veröffentlicht: (2025)
Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
von: Liu, Jiahao, et al.
Veröffentlicht: (2024)
von: Liu, Jiahao, et al.
Veröffentlicht: (2024)
Speculative Decoding with a Speculative Vocabulary
von: Williams, Miles, et al.
Veröffentlicht: (2026)
von: Williams, Miles, et al.
Veröffentlicht: (2026)
Nearest Neighbor Speculative Decoding for LLM Generation and Attribution
von: Li, Minghan, et al.
Veröffentlicht: (2024)
von: Li, Minghan, et al.
Veröffentlicht: (2024)
SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
von: Svirschevski, Ruslan, et al.
Veröffentlicht: (2024)
von: Svirschevski, Ruslan, et al.
Veröffentlicht: (2024)
Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inference
von: Zhang, Libo, et al.
Veröffentlicht: (2024)
von: Zhang, Libo, et al.
Veröffentlicht: (2024)
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
von: Shi, Luohe, et al.
Veröffentlicht: (2025)
von: Shi, Luohe, et al.
Veröffentlicht: (2025)
See the Forest for the Trees: Loosely Speculative Decoding via Visual-Semantic Guidance for Efficient Inference of Video LLMs
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
von: Ji, Yicheng, et al.
Veröffentlicht: (2026)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
von: Tang, Bangsheng, et al.
Veröffentlicht: (2025)
von: Tang, Bangsheng, et al.
Veröffentlicht: (2025)
Speculative Decoding and Beyond: An In-Depth Survey of Techniques
von: Hu, Yunhai, et al.
Veröffentlicht: (2025)
von: Hu, Yunhai, et al.
Veröffentlicht: (2025)
Beyond Tokens: Semantic-Aware Speculative Decoding for Efficient Inference by Probing Internal States
von: Dong, Ximing, et al.
Veröffentlicht: (2026)
von: Dong, Ximing, et al.
Veröffentlicht: (2026)
SPEED: Speculative Pipelined Execution for Efficient Decoding
von: Hooper, Coleman, et al.
Veröffentlicht: (2023)
von: Hooper, Coleman, et al.
Veröffentlicht: (2023)
Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding
von: Ryu, Hyun, et al.
Veröffentlicht: (2024)
von: Ryu, Hyun, et al.
Veröffentlicht: (2024)
HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative Decoding
von: Liu, Siran, et al.
Veröffentlicht: (2025)
von: Liu, Siran, et al.
Veröffentlicht: (2025)
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
von: Sun, Shengyin, et al.
Veröffentlicht: (2025)
von: Sun, Shengyin, et al.
Veröffentlicht: (2025)
Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional Verification
von: Wang, Jikai, et al.
Veröffentlicht: (2025)
von: Wang, Jikai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration
von: Xia, Heming, et al.
Veröffentlicht: (2024) -
ToolSpec: Accelerating Tool Calling via Schema-Aware and Retrieval-Augmented Speculative Decoding
von: Xia, Heming, et al.
Veröffentlicht: (2026) -
Merlin's Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting
von: Xia, Heming, et al.
Veröffentlicht: (2025) -
Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
von: Xia, Heming, et al.
Veröffentlicht: (2024) -
GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative Decoding
von: Du, Cunxiao, et al.
Veröffentlicht: (2024)