Batch Speculative Decoding Done Right
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Ranran Haoran, Dey, Soumik, Mishra, Ashirbad, Wu, Hansi, Li, Binbin, Zhang, Rui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Lazy to Prolific: Tackling Missing Labels in Open Vocabulary Extreme Classification by Positive-Unlabeled Sequence Learning
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2024)
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2024)
BroadGen: A Framework for Generating Effective and Efficient Advertiser Broad Match Keyphrase Recommendations
von: Mishra, Ashirbad, et al.
Veröffentlicht: (2025)
von: Mishra, Ashirbad, et al.
Veröffentlicht: (2025)
To Judge or not to Judge: Using LLM Judgements for Advertiser Keyphrase Relevance at eBay
von: Dey, Soumik, et al.
Veröffentlicht: (2025)
von: Dey, Soumik, et al.
Veröffentlicht: (2025)
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
MineDraft: A Framework for Batch Parallel Speculative Decoding
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
RASD: Retrieval-Augmented Speculative Decoding
von: Quan, Guofeng, et al.
Veröffentlicht: (2025)
von: Quan, Guofeng, et al.
Veröffentlicht: (2025)
Banking Done Right: Redefining Retail Banking with Language-Centric AI
von: Chua, Xin Jie, et al.
Veröffentlicht: (2025)
von: Chua, Xin Jie, et al.
Veröffentlicht: (2025)
SAM Decoding: Speculative Decoding via Suffix Automaton
von: Hu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2024)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
PACER: Blockwise Pre-verification for Speculative Decoding with Adaptive Length
von: Zhang, Situo, et al.
Veröffentlicht: (2026)
von: Zhang, Situo, et al.
Veröffentlicht: (2026)
The Disparate Impacts of Speculative Decoding
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
Cross-Attention Speculative Decoding
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
Speculative Decoding: Performance or Illusion?
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
Mamba Drafters for Speculative Decoding
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
Constrained Decoding with Speculative Lookaheads
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
Entropy-Aware Speculative Decoding Toward Improved LLM Reasoning
von: Su, Tiancheng, et al.
Veröffentlicht: (2025)
von: Su, Tiancheng, et al.
Veröffentlicht: (2025)
Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding
von: Jin, Tao, et al.
Veröffentlicht: (2026)
von: Jin, Tao, et al.
Veröffentlicht: (2026)
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
LLMDistill4Ads: Using Cross-Encoders to Distill from LLM Signals for Advertiser Keyphrase Recommendations at eBay
von: Dey, Soumik, et al.
Veröffentlicht: (2025)
von: Dey, Soumik, et al.
Veröffentlicht: (2025)
Dr. Kernel: Reinforcement Learning Done Right for Triton Kernel Generations
von: Liu, Wei, et al.
Veröffentlicht: (2026)
von: Liu, Wei, et al.
Veröffentlicht: (2026)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
Cacheback: Speculative Decoding With Nothing But Cache
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding
von: Su, Xin, et al.
Veröffentlicht: (2026)
von: Su, Xin, et al.
Veröffentlicht: (2026)
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
von: Li, Ruanjun, et al.
Veröffentlicht: (2025)
von: Li, Ruanjun, et al.
Veröffentlicht: (2025)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding
von: Song, Feifan, et al.
Veröffentlicht: (2025)
von: Song, Feifan, et al.
Veröffentlicht: (2025)
PASemiQA: Plan-Assisted Agent for Question Answering on Semi-Structured Data with Text and Relational Information
von: Yang, Hansi, et al.
Veröffentlicht: (2025)
von: Yang, Hansi, et al.
Veröffentlicht: (2025)
C2T: A Classifier-Based Tree Construction Method in Speculative Decoding
von: Huo, Feiye, et al.
Veröffentlicht: (2025)
von: Huo, Feiye, et al.
Veröffentlicht: (2025)
Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
GPT-3 Powered Information Extraction for Building Robust Knowledge Bases
von: Choudhury, Ritabrata Roy, et al.
Veröffentlicht: (2024)
von: Choudhury, Ritabrata Roy, et al.
Veröffentlicht: (2024)
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
von: Li, Guanghao, et al.
Veröffentlicht: (2025)
Traversal Verification for Speculative Tree Decoding
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
von: Chen, Yingfa, et al.
Veröffentlicht: (2026)
von: Chen, Yingfa, et al.
Veröffentlicht: (2026)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
Are We Done with MMLU?
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2024)
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
From Lazy to Prolific: Tackling Missing Labels in Open Vocabulary Extreme Classification by Positive-Unlabeled Sequence Learning
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2024) -
BroadGen: A Framework for Generating Effective and Efficient Advertiser Broad Match Keyphrase Recommendations
von: Mishra, Ashirbad, et al.
Veröffentlicht: (2025) -
To Judge or not to Judge: Using LLM Judgements for Advertiser Keyphrase Relevance at eBay
von: Dey, Soumik, et al.
Veröffentlicht: (2025) -
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025) -
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)