FastEagle: Cascaded Drafting for Accelerating Speculative Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Haiduo, Song, Jiangcheng, Zhao, Wenzhe, Ren, Pengju |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation Trainer
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
NMS: Efficient Edge DNN Training via Near-Memory Sampling on Manifolds
von: Zhao, Boran, et al.
Veröffentlicht: (2025)
von: Zhao, Boran, et al.
Veröffentlicht: (2025)
SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Cascade Speculative Drafting for Even Faster LLM Inference
von: Chen, Ziyi, et al.
Veröffentlicht: (2023)
von: Chen, Ziyi, et al.
Veröffentlicht: (2023)
KernelDNA: Dynamic Kernel Sharing via Decoupled Naive Adapters
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
POSS: Position Specialist Generates Better Draft for Speculative Decoding
von: Huang, Langlin, et al.
Veröffentlicht: (2025)
von: Huang, Langlin, et al.
Veröffentlicht: (2025)
SpecVLM: Fast Speculative Decoding in Vision-Language Models
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
When Drafts Evolve: Speculative Decoding Meets Online Learning
von: Qian, Yu-Yang, et al.
Veröffentlicht: (2026)
von: Qian, Yu-Yang, et al.
Veröffentlicht: (2026)
FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
Draft, Verify, and Improve: Toward Training-Aware Speculative Decoding
von: Bhansali, Shrenik, et al.
Veröffentlicht: (2025)
von: Bhansali, Shrenik, et al.
Veröffentlicht: (2025)
SparseMap: A Sparse Tensor Accelerator Framework Based on Evolution Strategy
von: Zhao, Boran, et al.
Veröffentlicht: (2025)
von: Zhao, Boran, et al.
Veröffentlicht: (2025)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2025)
von: Li, Jinze, et al.
Veröffentlicht: (2025)
DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation
von: Liu, Zining, et al.
Veröffentlicht: (2026)
von: Liu, Zining, et al.
Veröffentlicht: (2026)
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
von: Shen, Yuhao, et al.
Veröffentlicht: (2026)
von: Shen, Yuhao, et al.
Veröffentlicht: (2026)
Partial Channel Network: Compute Fewer, Perform Better
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
von: Pan, Rui, et al.
Veröffentlicht: (2025)
von: Pan, Rui, et al.
Veröffentlicht: (2025)
Faster Cascades via Speculative Decoding
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
TABED: Test-Time Adaptive Ensemble Drafting for Robust Speculative Decoding in LVLMs
von: Lee, Minjae, et al.
Veröffentlicht: (2026)
von: Lee, Minjae, et al.
Veröffentlicht: (2026)
ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
von: Hong, Fenglu, et al.
Veröffentlicht: (2025)
von: Hong, Fenglu, et al.
Veröffentlicht: (2025)
Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding
von: Shoham, Ofir Ben
Veröffentlicht: (2026)
von: Shoham, Ofir Ben
Veröffentlicht: (2026)
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
von: Ning, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Ning, Zhiyuan, et al.
Veröffentlicht: (2025)
Fast Inference via Hierarchical Speculative Decoding
von: Mohri, Clara, et al.
Veröffentlicht: (2025)
von: Mohri, Clara, et al.
Veröffentlicht: (2025)
OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding
von: Ramakrishnan, Ramchalam Kinattinkara, et al.
Veröffentlicht: (2025)
von: Ramakrishnan, Ramchalam Kinattinkara, et al.
Veröffentlicht: (2025)
BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding
von: He, Liang, et al.
Veröffentlicht: (2026)
von: He, Liang, et al.
Veröffentlicht: (2026)
MineDraft: A Framework for Batch Parallel Speculative Decoding
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
von: Goel, Raghavv, et al.
Veröffentlicht: (2024)
von: Goel, Raghavv, et al.
Veröffentlicht: (2024)
GeGS-PCR: Effective and Robust 3D Point Cloud Registration with Two-Stage Color-Enhanced Geometric-3DGS Fusion
von: Tian, Jiayi, et al.
Veröffentlicht: (2026)
von: Tian, Jiayi, et al.
Veröffentlicht: (2026)
Speculative Speculative Decoding
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
Accelerating Time Series Foundation Models with Speculative Decoding
von: Subbaraman, Pranav, et al.
Veröffentlicht: (2025)
von: Subbaraman, Pranav, et al.
Veröffentlicht: (2025)
CATS: Cascaded Adaptive Tree Speculation for Memory-Limited LLM Inference Acceleration
von: Han, Yuning, et al.
Veröffentlicht: (2026)
von: Han, Yuning, et al.
Veröffentlicht: (2026)
Make Every Draft Count: Hidden State based Speculative Decoding
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
von: Zhao, Yilong, et al.
Veröffentlicht: (2025)
Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decoding
von: Sun, Shuoyang, et al.
Veröffentlicht: (2026)
von: Sun, Shuoyang, et al.
Veröffentlicht: (2026)
Nearly Lossless Adaptive Bit Switching
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
von: Huang, Haiduo, et al.
Veröffentlicht: (2025)
TPP-SD: Accelerating Transformer Point Process Sampling with Speculative Decoding
von: Gong, Shukai, et al.
Veröffentlicht: (2025)
von: Gong, Shukai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation Trainer
von: Huang, Haiduo, et al.
Veröffentlicht: (2025) -
NMS: Efficient Edge DNN Training via Near-Memory Sampling on Manifolds
von: Zhao, Boran, et al.
Veröffentlicht: (2025) -
SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
von: Huang, Haiduo, et al.
Veröffentlicht: (2025) -
Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE
von: Huang, Haiduo, et al.
Veröffentlicht: (2025) -
Cascade Speculative Drafting for Even Faster LLM Inference
von: Chen, Ziyi, et al.
Veröffentlicht: (2023)