Temperature-Centric Investigation of Speculative Decoding with Knowledge Distillation
Fuente:
arXiv
Salvato in:
| Autori principali: | Ouyang, Siru, Wang, Shuohang, Jiang, Minhao, Zhong, Ming, Yu, Donghan, Han, Jiawei, Shen, Yelong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multi-LoRA Composition for Image Generation
di: Zhong, Ming, et al.
Pubblicazione: (2024)
di: Zhong, Ming, et al.
Pubblicazione: (2024)
Investigating Data Contamination for Pre-training Language Models
di: Jiang, Minhao, et al.
Pubblicazione: (2024)
di: Jiang, Minhao, et al.
Pubblicazione: (2024)
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
di: Xiao, Zilin, et al.
Pubblicazione: (2024)
di: Xiao, Zilin, et al.
Pubblicazione: (2024)
A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
di: Jiang, Pengcheng, et al.
Pubblicazione: (2025)
di: Jiang, Pengcheng, et al.
Pubblicazione: (2025)
Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning
di: Chen, Yifang, et al.
Pubblicazione: (2024)
di: Chen, Yifang, et al.
Pubblicazione: (2024)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
di: Zhou, Yongchao, et al.
Pubblicazione: (2023)
di: Zhou, Yongchao, et al.
Pubblicazione: (2023)
AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
di: Hu, Yuezhou, et al.
Pubblicazione: (2025)
di: Hu, Yuezhou, et al.
Pubblicazione: (2025)
Synergizing Unsupervised Episode Detection with LLMs for Large-Scale News Events
di: Kargupta, Priyanka, et al.
Pubblicazione: (2024)
di: Kargupta, Priyanka, et al.
Pubblicazione: (2024)
RAS: Retrieval-And-Structuring for Knowledge-Intensive LLM Generation
di: Jiang, Pengcheng, et al.
Pubblicazione: (2025)
di: Jiang, Pengcheng, et al.
Pubblicazione: (2025)
OntoType: Ontology-Guided and Pre-Trained Language Model Assisted Fine-Grained Entity Typing
di: Komarlu, Tanay, et al.
Pubblicazione: (2023)
di: Komarlu, Tanay, et al.
Pubblicazione: (2023)
Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
di: Ren, Liliang, et al.
Pubblicazione: (2025)
di: Ren, Liliang, et al.
Pubblicazione: (2025)
Cross-Attention Speculative Decoding
di: Zhong, Wei, et al.
Pubblicazione: (2025)
di: Zhong, Wei, et al.
Pubblicazione: (2025)
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
di: Zhang, Zhen, et al.
Pubblicazione: (2025)
di: Zhang, Zhen, et al.
Pubblicazione: (2025)
LoRC: Low-Rank Compression for LLMs KV Cache with a Progressive Compression Strategy
di: Zhang, Rongzhi, et al.
Pubblicazione: (2024)
di: Zhang, Rongzhi, et al.
Pubblicazione: (2024)
Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling
di: Xu, Wenda, et al.
Pubblicazione: (2024)
di: Xu, Wenda, et al.
Pubblicazione: (2024)
Cacheback: Speculative Decoding With Nothing But Cache
di: Ma, Zhiyao, et al.
Pubblicazione: (2025)
di: Ma, Zhiyao, et al.
Pubblicazione: (2025)
Performance-Driven Policy Optimization for Speculative Decoding with Adaptive Windowing
di: Jiang, Jie, et al.
Pubblicazione: (2026)
di: Jiang, Jie, et al.
Pubblicazione: (2026)
Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval
di: Jiang, Pengcheng, et al.
Pubblicazione: (2024)
di: Jiang, Pengcheng, et al.
Pubblicazione: (2024)
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
di: Wang, Yiping, et al.
Pubblicazione: (2024)
di: Wang, Yiping, et al.
Pubblicazione: (2024)
Benchmarking Retrieval-Augmented Generation for Chemistry
di: Zhong, Xianrui, et al.
Pubblicazione: (2025)
di: Zhong, Xianrui, et al.
Pubblicazione: (2025)
CLaSp: In-Context Layer Skip for Self-Speculative Decoding
di: Chen, Longze, et al.
Pubblicazione: (2025)
di: Chen, Longze, et al.
Pubblicazione: (2025)
Speculative Decoding with a Speculative Vocabulary
di: Williams, Miles, et al.
Pubblicazione: (2026)
di: Williams, Miles, et al.
Pubblicazione: (2026)
Graph-Structured Speculative Decoding
di: Gong, Zhuocheng, et al.
Pubblicazione: (2024)
di: Gong, Zhuocheng, et al.
Pubblicazione: (2024)
Decoding Speculative Decoding
di: Yan, Minghao, et al.
Pubblicazione: (2024)
di: Yan, Minghao, et al.
Pubblicazione: (2024)
Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding
di: Wang, Ziyao, et al.
Pubblicazione: (2025)
di: Wang, Ziyao, et al.
Pubblicazione: (2025)
Speculative Pipeline Decoding: Higher-Accruacy and Zero-Bubble Speculation via Pipeline Parallelism
di: Yu, Yijiong, et al.
Pubblicazione: (2026)
di: Yu, Yijiong, et al.
Pubblicazione: (2026)
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
di: Brown, Oscar, et al.
Pubblicazione: (2024)
di: Brown, Oscar, et al.
Pubblicazione: (2024)
RASD: Retrieval-Augmented Speculative Decoding
di: Quan, Guofeng, et al.
Pubblicazione: (2025)
di: Quan, Guofeng, et al.
Pubblicazione: (2025)
Speculative Contrastive Decoding
di: Yuan, Hongyi, et al.
Pubblicazione: (2023)
di: Yuan, Hongyi, et al.
Pubblicazione: (2023)
Improving Multi-candidate Speculative Decoding
di: Lu, Xiaofan, et al.
Pubblicazione: (2024)
di: Lu, Xiaofan, et al.
Pubblicazione: (2024)
Scaling Laws for Speculative Decoding
di: Yan, Siyuan, et al.
Pubblicazione: (2025)
di: Yan, Siyuan, et al.
Pubblicazione: (2025)
Fast Best-of-N Decoding via Speculative Rejection
di: Sun, Hanshi, et al.
Pubblicazione: (2024)
di: Sun, Hanshi, et al.
Pubblicazione: (2024)
A Theoretical Perspective for Speculative Decoding Algorithm
di: Yin, Ming, et al.
Pubblicazione: (2024)
di: Yin, Ming, et al.
Pubblicazione: (2024)
GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative Decoding
di: Du, Cunxiao, et al.
Pubblicazione: (2024)
di: Du, Cunxiao, et al.
Pubblicazione: (2024)
Decoder-based Sense Knowledge Distillation
di: Wang, Qitong, et al.
Pubblicazione: (2026)
di: Wang, Qitong, et al.
Pubblicazione: (2026)
Speculative Decoding: Performance or Illusion?
di: Liu, Xiaoxuan, et al.
Pubblicazione: (2025)
di: Liu, Xiaoxuan, et al.
Pubblicazione: (2025)
Boosting Lossless Speculative Decoding via Feature Sampling and Partial Alignment Distillation
di: Gui, Lujun, et al.
Pubblicazione: (2024)
di: Gui, Lujun, et al.
Pubblicazione: (2024)
RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding
di: Hoshino, Yuichiro, et al.
Pubblicazione: (2025)
di: Hoshino, Yuichiro, et al.
Pubblicazione: (2025)
Mamba Drafters for Speculative Decoding
di: Choi, Daewon, et al.
Pubblicazione: (2025)
di: Choi, Daewon, et al.
Pubblicazione: (2025)
Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning
di: Zhang, Jiebin, et al.
Pubblicazione: (2026)
di: Zhang, Jiebin, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Multi-LoRA Composition for Image Generation
di: Zhong, Ming, et al.
Pubblicazione: (2024) -
Investigating Data Contamination for Pre-training Language Models
di: Jiang, Minhao, et al.
Pubblicazione: (2024) -
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
di: Xiao, Zilin, et al.
Pubblicazione: (2024) -
A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
di: Jiang, Pengcheng, et al.
Pubblicazione: (2025) -
Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning
di: Chen, Yifang, et al.
Pubblicazione: (2024)