DART: Diffusion-Inspired Speculative Decoding for Fast LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Fuliang, Li, Xue, Zhao, Ketai, Gao, Yinxi, Zhou, Ziyan, Zhang, Zhonghui, Wang, Zhibin, Dou, Wanchun, Zhong, Sheng, Tian, Chen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters
von: Yi, Euiin, et al.
Veröffentlicht: (2024)
von: Yi, Euiin, et al.
Veröffentlicht: (2024)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
Tutorial Proposal: Speculative Decoding for Efficient LLM Inference
von: Xia, Heming, et al.
Veröffentlicht: (2025)
von: Xia, Heming, et al.
Veröffentlicht: (2025)
Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration
von: Xia, Heming, et al.
Veröffentlicht: (2024)
von: Xia, Heming, et al.
Veröffentlicht: (2024)
S2D2: Fast Decoding for Diffusion LLMs via Training-Free Self-Speculation
von: Han, Ligong, et al.
Veröffentlicht: (2026)
von: Han, Ligong, et al.
Veröffentlicht: (2026)
SSV: Sparse Speculative Verification for Efficient LLM Inference
von: Wang, Zhibin, et al.
Veröffentlicht: (2026)
von: Wang, Zhibin, et al.
Veröffentlicht: (2026)
SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
von: Svirschevski, Ruslan, et al.
Veröffentlicht: (2024)
von: Svirschevski, Ruslan, et al.
Veröffentlicht: (2024)
DFlash: Block Diffusion for Flash Speculative Decoding
von: Chen, Jian, et al.
Veröffentlicht: (2026)
von: Chen, Jian, et al.
Veröffentlicht: (2026)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
SDSAT: Accelerating LLM Inference through Speculative Decoding with Semantic Adaptive Tokens
von: Liu, Chengbo, et al.
Veröffentlicht: (2024)
von: Liu, Chengbo, et al.
Veröffentlicht: (2024)
Self Speculative Decoding for Diffusion Large Language Models
von: Gao, Yifeng, et al.
Veröffentlicht: (2025)
von: Gao, Yifeng, et al.
Veröffentlicht: (2025)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
Speculative Streaming: Fast LLM Inference without Auxiliary Models
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2024)
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2024)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
Nearest Neighbor Speculative Decoding for LLM Generation and Attribution
von: Li, Minghan, et al.
Veröffentlicht: (2024)
von: Li, Minghan, et al.
Veröffentlicht: (2024)
RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding
von: Chen, Guanzheng, et al.
Veröffentlicht: (2025)
von: Chen, Guanzheng, et al.
Veröffentlicht: (2025)
Automatic Task Detection and Heterogeneous LLM Speculative Decoding
von: Ge, Danying, et al.
Veröffentlicht: (2025)
von: Ge, Danying, et al.
Veröffentlicht: (2025)
Fast Best-of-N Decoding via Speculative Rejection
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
Speculative Contrastive Decoding
von: Yuan, Hongyi, et al.
Veröffentlicht: (2023)
von: Yuan, Hongyi, et al.
Veröffentlicht: (2023)
Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inference
von: Zhang, Libo, et al.
Veröffentlicht: (2024)
von: Zhang, Libo, et al.
Veröffentlicht: (2024)
Cacheback: Speculative Decoding With Nothing But Cache
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism
von: Liu, Jiahao, et al.
Veröffentlicht: (2024)
von: Liu, Jiahao, et al.
Veröffentlicht: (2024)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
Cross-Attention Speculative Decoding
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
DReSD: Dense Retrieval for Speculative Decoding
von: Gritta, Milan, et al.
Veröffentlicht: (2025)
von: Gritta, Milan, et al.
Veröffentlicht: (2025)
KOALA: Enhancing Speculative Decoding for LLM via Multi-Layer Draft Heads with Adversarial Learning
von: Zhang, Kaiqi, et al.
Veröffentlicht: (2024)
von: Zhang, Kaiqi, et al.
Veröffentlicht: (2024)
TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar
von: Li, Yinxi, et al.
Veröffentlicht: (2025)
von: Li, Yinxi, et al.
Veröffentlicht: (2025)
SpecTr-GBV: Multi-Draft Block Verification Accelerating Speculative Decoding
von: Lin, Yijun, et al.
Veröffentlicht: (2026)
von: Lin, Yijun, et al.
Veröffentlicht: (2026)
Clover-2: Accurate Inference for Regressive Lightweight Speculative Decoding
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference
von: Zhou, Xuwen, et al.
Veröffentlicht: (2026)
von: Zhou, Xuwen, et al.
Veröffentlicht: (2026)
Fast Large Language Model Collaborative Decoding via Speculation
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
von: Fu, Jiale, et al.
Veröffentlicht: (2025)
Cost-Aware Diffusion Draft Trees for Speculative Decoding
von: Zhang, Shuai, et al.
Veröffentlicht: (2026)
von: Zhang, Shuai, et al.
Veröffentlicht: (2026)
Accelerating Speculative Decoding with Block Diffusion Draft Trees
von: Ringel, Liran, et al.
Veröffentlicht: (2026)
von: Ringel, Liran, et al.
Veröffentlicht: (2026)
Speculative Decoding with a Speculative Vocabulary
von: Williams, Miles, et al.
Veröffentlicht: (2026)
von: Williams, Miles, et al.
Veröffentlicht: (2026)
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding
von: Sun, Shuoyang, et al.
Veröffentlicht: (2026)
von: Sun, Shuoyang, et al.
Veröffentlicht: (2026)
Factorization-Error-Free Discrete Diffusion Language Model via Speculative Decoding
von: Fang, Xun, et al.
Veröffentlicht: (2026)
von: Fang, Xun, et al.
Veröffentlicht: (2026)
Decoding Speculative Decoding
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters
von: Yi, Euiin, et al.
Veröffentlicht: (2024) -
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025) -
Tutorial Proposal: Speculative Decoding for Efficient LLM Inference
von: Xia, Heming, et al.
Veröffentlicht: (2025) -
Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025) -
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration
von: Xia, Heming, et al.
Veröffentlicht: (2024)