DistillSpec: Improving Speculative Decoding via Knowledge Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yongchao, Lyu, Kaifeng, Rawat, Ankit Singh, Menon, Aditya Krishna, Rostamizadeh, Afshin, Kumar, Sanjiv, Kagy, Jean-François, Agarwal, Rishabh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Faster Cascades via Speculative Decoding
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
SoftSRV: Learn to Generate Targeted Synthetic Data
von: DeSalvo, Giulia, et al.
Veröffentlicht: (2024)
von: DeSalvo, Giulia, et al.
Veröffentlicht: (2024)
SpacTor-T5: Pre-training T5 Models with Span Corruption and Replaced Token Detection
von: Ye, Ke, et al.
Veröffentlicht: (2024)
von: Ye, Ke, et al.
Veröffentlicht: (2024)
Think before you speak: Training Language Models With Pause Tokens
von: Goyal, Sachin, et al.
Veröffentlicht: (2023)
von: Goyal, Sachin, et al.
Veröffentlicht: (2023)
When Does Confidence-Based Cascade Deferral Suffice?
von: Jitkrittum, Wittawat, et al.
Veröffentlicht: (2023)
von: Jitkrittum, Wittawat, et al.
Veröffentlicht: (2023)
Language Model Cascades: Token-level uncertainty and beyond
von: Gupta, Neha, et al.
Veröffentlicht: (2024)
von: Gupta, Neha, et al.
Veröffentlicht: (2024)
HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative Decoding
von: Liu, Siran, et al.
Veröffentlicht: (2025)
von: Liu, Siran, et al.
Veröffentlicht: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs
von: Rawat, Ankit Singh, et al.
Veröffentlicht: (2024)
von: Rawat, Ankit Singh, et al.
Veröffentlicht: (2024)
Analyzing Similarity Metrics for Data Selection for Language Model Pretraining
von: Sam, Dylan, et al.
Veröffentlicht: (2025)
von: Sam, Dylan, et al.
Veröffentlicht: (2025)
Temperature-Centric Investigation of Speculative Decoding with Knowledge Distillation
von: Ouyang, Siru, et al.
Veröffentlicht: (2024)
von: Ouyang, Siru, et al.
Veröffentlicht: (2024)
HiSpec: Hierarchical Speculative Decoding for LLMs
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
von: Hu, Yuezhou, et al.
Veröffentlicht: (2025)
von: Hu, Yuezhou, et al.
Veröffentlicht: (2025)
Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling
von: Xu, Wenda, et al.
Veröffentlicht: (2024)
von: Xu, Wenda, et al.
Veröffentlicht: (2024)
Algorithms for Learning Kernels Based on Centered Alignment
von: Cortes, Corinna, et al.
Veröffentlicht: (2012)
von: Cortes, Corinna, et al.
Veröffentlicht: (2012)
Efficient Document Ranking with Learnable Late Interactions
von: Ji, Ziwei, et al.
Veröffentlicht: (2024)
von: Ji, Ziwei, et al.
Veröffentlicht: (2024)
Simple Unsupervised Knowledge Distillation With Space Similarity
von: Singh, Aditya, et al.
Veröffentlicht: (2024)
von: Singh, Aditya, et al.
Veröffentlicht: (2024)
SpecMemo: Speculative Decoding is in Your Pocket
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
von: Agarwal, Rishabh, et al.
Veröffentlicht: (2023)
von: Agarwal, Rishabh, et al.
Veröffentlicht: (2023)
Don't Throw Away Data: Better Sequence Knowledge Distillation
von: Wang, Jun, et al.
Veröffentlicht: (2024)
von: Wang, Jun, et al.
Veröffentlicht: (2024)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
Decoding Speculative Decoding
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
Regression-aware Inference with LLMs
von: Lukasik, Michal, et al.
Veröffentlicht: (2024)
von: Lukasik, Michal, et al.
Veröffentlicht: (2024)
On student-teacher deviations in distillation: does it pay to disobey?
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2023)
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2023)
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation
von: Phan, Phuc, et al.
Veröffentlicht: (2024)
von: Phan, Phuc, et al.
Veröffentlicht: (2024)
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
SpecMamba: Accelerating Mamba Inference on FPGA with Speculative Decoding
von: Zhong, Linfeng, et al.
Veröffentlicht: (2025)
von: Zhong, Linfeng, et al.
Veröffentlicht: (2025)
SpecTr: Fast Speculative Decoding via Optimal Transport
von: Sun, Ziteng, et al.
Veröffentlicht: (2023)
von: Sun, Ziteng, et al.
Veröffentlicht: (2023)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
von: Hou, Yunlong, et al.
Veröffentlicht: (2025)
von: Hou, Yunlong, et al.
Veröffentlicht: (2025)
MoE-Spec: Expert Budgeting for Efficient Speculative Decoding
von: McDanel, Bradley, et al.
Veröffentlicht: (2026)
von: McDanel, Bradley, et al.
Veröffentlicht: (2026)
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
von: Sun, Ryan, et al.
Veröffentlicht: (2024)
von: Sun, Ryan, et al.
Veröffentlicht: (2024)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
SpecPV: Improving Self-Speculative Decoding for Long-Context Generation via Partial Verification
von: Tan, Zhendong, et al.
Veröffentlicht: (2025)
von: Tan, Zhendong, et al.
Veröffentlicht: (2025)
ChromaDistill: Colorizing Monochrome Radiance Fields with Knowledge Distillation
von: Dhiman, Ankit, et al.
Veröffentlicht: (2023)
von: Dhiman, Ankit, et al.
Veröffentlicht: (2023)
LogitSpec: Accelerating Retrieval-based Speculative Decoding via Next Next Token Speculation
von: Liu, Tianyu, et al.
Veröffentlicht: (2025)
von: Liu, Tianyu, et al.
Veröffentlicht: (2025)
Decoder-based Sense Knowledge Distillation
von: Wang, Qitong, et al.
Veröffentlicht: (2026)
von: Wang, Qitong, et al.
Veröffentlicht: (2026)
Low-Resolution Chest X-ray Classification via Knowledge Distillation and Multi-task Learning
von: Akhter, Yasmeena, et al.
Veröffentlicht: (2024)
von: Akhter, Yasmeena, et al.
Veröffentlicht: (2024)
BrainDistill: Implantable Motor Decoding with Task-Specific Knowledge Distillation
von: Xie, Yuhan, et al.
Veröffentlicht: (2026)
von: Xie, Yuhan, et al.
Veröffentlicht: (2026)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Faster Cascades via Speculative Decoding
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024) -
SoftSRV: Learn to Generate Targeted Synthetic Data
von: DeSalvo, Giulia, et al.
Veröffentlicht: (2024) -
SpacTor-T5: Pre-training T5 Models with Span Corruption and Replaced Token Detection
von: Ye, Ke, et al.
Veröffentlicht: (2024) -
Think before you speak: Training Language Models With Pause Tokens
von: Goyal, Sachin, et al.
Veröffentlicht: (2023) -
When Does Confidence-Based Cascade Deferral Suffice?
von: Jitkrittum, Wittawat, et al.
Veröffentlicht: (2023)