Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Fenglu, Raju, Ravi, Li, Jonathan Lingjie, Li, Bo, Thakker, Urmish, Ravichandran, Avinash, Jain, Swayambhoo, Hu, Changran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge
von: Raju, Ravi, et al.
Veröffentlicht: (2024)
von: Raju, Ravi, et al.
Veröffentlicht: (2024)
Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Models
von: Upasani, Shubhangi, et al.
Veröffentlicht: (2026)
von: Upasani, Shubhangi, et al.
Veröffentlicht: (2026)
Synthetic Document Question Answering in Hungarian
von: Li, Jonathan, et al.
Veröffentlicht: (2025)
von: Li, Jonathan, et al.
Veröffentlicht: (2025)
Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls
von: Upasani, Shubhangi, et al.
Veröffentlicht: (2026)
von: Upasani, Shubhangi, et al.
Veröffentlicht: (2026)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)
SubgoalXL: Subgoal-based Expert Learning for Theorem Proving
von: Zhao, Xueliang, et al.
Veröffentlicht: (2024)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2024)
Composition of Experts: A Modular Compound AI System Leveraging Large Language Models
von: Jain, Swayambhoo, et al.
Veröffentlicht: (2024)
von: Jain, Swayambhoo, et al.
Veröffentlicht: (2024)
SambaLingo: Teaching Large Language Models New Languages
von: Csaki, Zoltan, et al.
Veröffentlicht: (2024)
von: Csaki, Zoltan, et al.
Veröffentlicht: (2024)
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
von: Zhang, Qizheng, et al.
Veröffentlicht: (2025)
von: Zhang, Qizheng, et al.
Veröffentlicht: (2025)
The Limits of Long-Context Reasoning in Automated Bug Fixing
von: Raju, Ravi, et al.
Veröffentlicht: (2026)
von: Raju, Ravi, et al.
Veröffentlicht: (2026)
PEARL: Parallel Speculative Decoding with Adaptive Draft Length
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
von: Liu, Tianyu, et al.
Veröffentlicht: (2024)
Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning
von: Zhang, Jiebin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiebin, et al.
Veröffentlicht: (2026)
Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match
von: Li, Jinze, et al.
Veröffentlicht: (2025)
von: Li, Jinze, et al.
Veröffentlicht: (2025)
OPT-Tree: Speculative Decoding with Adaptive Draft Tree Structure
von: Wang, Jikai, et al.
Veröffentlicht: (2024)
von: Wang, Jikai, et al.
Veröffentlicht: (2024)
DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding
von: Zhang, Jiebin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiebin, et al.
Veröffentlicht: (2026)
Cost-Aware Diffusion Draft Trees for Speculative Decoding
von: Zhang, Shuai, et al.
Veröffentlicht: (2026)
von: Zhang, Shuai, et al.
Veröffentlicht: (2026)
Accelerating Speculative Decoding with Block Diffusion Draft Trees
von: Ringel, Liran, et al.
Veröffentlicht: (2026)
von: Ringel, Liran, et al.
Veröffentlicht: (2026)
Make Every Draft Count: Hidden State based Speculative Decoding
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
LongAttnComp: Cross-Family Context Compression for Long-Context Reasoning
von: Ji, Mengmeng, et al.
Veröffentlicht: (2026)
von: Ji, Mengmeng, et al.
Veröffentlicht: (2026)
Speculative Decoding for Verilog: Speed and Quality, All in One
von: Xu, Changran, et al.
Veröffentlicht: (2025)
von: Xu, Changran, et al.
Veröffentlicht: (2025)
Draft-OPD: On-Policy Distillation for Speculative Draft Models
von: Lei, Haodi, et al.
Veröffentlicht: (2026)
von: Lei, Haodi, et al.
Veröffentlicht: (2026)
DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding
von: Hu, Yunhai, et al.
Veröffentlicht: (2025)
von: Hu, Yunhai, et al.
Veröffentlicht: (2025)
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
von: Sun, Ryan, et al.
Veröffentlicht: (2024)
von: Sun, Ryan, et al.
Veröffentlicht: (2024)
DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting
von: Lv, Kai, et al.
Veröffentlicht: (2025)
von: Lv, Kai, et al.
Veröffentlicht: (2025)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
von: Zhang, Jun, et al.
Veröffentlicht: (2023)
von: Zhang, Jun, et al.
Veröffentlicht: (2023)
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
von: Zhao, Weilin, et al.
Veröffentlicht: (2024)
von: Zhao, Weilin, et al.
Veröffentlicht: (2024)
SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting
von: Shi, Weijie, et al.
Veröffentlicht: (2026)
von: Shi, Weijie, et al.
Veröffentlicht: (2026)
Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
von: Huang, Jianuo, et al.
Veröffentlicht: (2026)
von: Huang, Jianuo, et al.
Veröffentlicht: (2026)
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
von: Metel, Michael R., et al.
Veröffentlicht: (2024)
von: Metel, Michael R., et al.
Veröffentlicht: (2024)
PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding
von: An, Zihao, et al.
Veröffentlicht: (2026)
von: An, Zihao, et al.
Veröffentlicht: (2026)
POSS: Position Specialist Generates Better Draft for Speculative Decoding
von: Huang, Langlin, et al.
Veröffentlicht: (2025)
von: Huang, Langlin, et al.
Veröffentlicht: (2025)
Flatter Tokens are More Valuable for Speculative Draft Model Training
von: Fan, Jiaming, et al.
Veröffentlicht: (2026)
von: Fan, Jiaming, et al.
Veröffentlicht: (2026)
SpecTr-GBV: Multi-Draft Block Verification Accelerating Speculative Decoding
von: Lin, Yijun, et al.
Veröffentlicht: (2026)
von: Lin, Yijun, et al.
Veröffentlicht: (2026)
Fast Best-of-N Decoding via Speculative Rejection
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
HiSpec: Hierarchical Speculative Decoding for LLMs
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
Draft-Conditioned Constrained Decoding for Structured Generation in LLMs
von: Reddy, Avinash, et al.
Veröffentlicht: (2026)
von: Reddy, Avinash, et al.
Veröffentlicht: (2026)
ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge
von: Raju, Ravi, et al.
Veröffentlicht: (2024) -
Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Models
von: Upasani, Shubhangi, et al.
Veröffentlicht: (2026) -
Synthetic Document Question Answering in Hungarian
von: Li, Jonathan, et al.
Veröffentlicht: (2025) -
Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls
von: Upasani, Shubhangi, et al.
Veröffentlicht: (2026) -
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
von: Wang, Guangtao, et al.
Veröffentlicht: (2025)