Gespeichert in:
| Hauptverfasser: | Xiong, Yunfan, Zhang, Ruoyu, Li, Yanzeng, Wu, Tianhao, Zou, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2410.11744 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
MoE-Spec: Expert Budgeting for Efficient Speculative Decoding
von: McDanel, Bradley, et al.
Veröffentlicht: (2026)
von: McDanel, Bradley, et al.
Veröffentlicht: (2026)
EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU Utilization
von: Wu, Yize, et al.
Veröffentlicht: (2025)
von: Wu, Yize, et al.
Veröffentlicht: (2025)
SpecPipe: Accelerating Pipeline Parallelism-based LLM Inference with Speculative Decoding
von: Yin, Haofei, et al.
Veröffentlicht: (2025)
von: Yin, Haofei, et al.
Veröffentlicht: (2025)
HiSpec: Hierarchical Speculative Decoding for LLMs
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
TriSpec: Ternary Speculative Decoding via Lightweight Proxy Verification
von: Jiang, Haoyun, et al.
Veröffentlicht: (2026)
von: Jiang, Haoyun, et al.
Veröffentlicht: (2026)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
von: Hou, Yunlong, et al.
Veröffentlicht: (2025)
von: Hou, Yunlong, et al.
Veröffentlicht: (2025)
SpecMemo: Speculative Decoding is in Your Pocket
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
Faster Cascades via Speculative Decoding
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
von: Narasimhan, Harikrishna, et al.
Veröffentlicht: (2024)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding
von: Li, Shenggui, et al.
Veröffentlicht: (2026)
von: Li, Shenggui, et al.
Veröffentlicht: (2026)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
von: Zhou, Yongchao, et al.
Veröffentlicht: (2023)
SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
SpecMER: Fast Protein Generation with K-mer Guided Speculative Decoding
von: Walton, Thomas, et al.
Veröffentlicht: (2025)
von: Walton, Thomas, et al.
Veröffentlicht: (2025)
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
von: Ning, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Ning, Zhiyuan, et al.
Veröffentlicht: (2025)
Annealed Relaxation of Speculative Decoding for Faster Autoregressive Image Generation
von: Li, Xingyao, et al.
Veröffentlicht: (2026)
von: Li, Xingyao, et al.
Veröffentlicht: (2026)
SpecPV: Improving Self-Speculative Decoding for Long-Context Generation via Partial Verification
von: Tan, Zhendong, et al.
Veröffentlicht: (2025)
von: Tan, Zhendong, et al.
Veröffentlicht: (2025)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths
von: Huang, Kaixuan, et al.
Veröffentlicht: (2024)
von: Huang, Kaixuan, et al.
Veröffentlicht: (2024)
ML-SpecQD: Multi-Level Speculative Decoding with Quantized Drafts
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
von: Georganas, Evangelos, et al.
Veröffentlicht: (2025)
SpecTr: Fast Speculative Decoding via Optimal Transport
von: Sun, Ziteng, et al.
Veröffentlicht: (2023)
von: Sun, Ziteng, et al.
Veröffentlicht: (2023)
Spec-VLA: Speculative Decoding for Vision-Language-Action Models with Relaxed Acceptance
von: Wang, Songsheng, et al.
Veröffentlicht: (2025)
von: Wang, Songsheng, et al.
Veröffentlicht: (2025)
SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
von: Plaksin, Anton, et al.
Veröffentlicht: (2026)
Self-Speculative Biased Decoding for Faster Re-Translation
von: Zeng, Linxiao, et al.
Veröffentlicht: (2025)
von: Zeng, Linxiao, et al.
Veröffentlicht: (2025)
Think Before You Accept: Semantic Reflective Verification for Faster Speculative Decoding
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
Greedy Multi-Path Block Verification for Faster Decoding in Speculative Sampling
von: Thomas, Rahul, et al.
Veröffentlicht: (2026)
von: Thomas, Rahul, et al.
Veröffentlicht: (2026)
SpecExtend: A Drop-in Enhancement for Speculative Decoding of Long Sequences
von: Cha, Jungyoub, et al.
Veröffentlicht: (2025)
von: Cha, Jungyoub, et al.
Veröffentlicht: (2025)
Speculative Speculative Decoding
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2026)
KnapSpec: Self-Speculative Decoding via Adaptive Layer Selection as a Knapsack Problem
von: Cha, Seongjin, et al.
Veröffentlicht: (2026)
von: Cha, Seongjin, et al.
Veröffentlicht: (2026)
SpecAttn: Speculating Sparse Attention
von: Shah, Harsh
Veröffentlicht: (2025)
von: Shah, Harsh
Veröffentlicht: (2025)
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
von: Shukla, Shikhar
Veröffentlicht: (2026)
von: Shukla, Shikhar
Veröffentlicht: (2026)
Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking
von: Li, Yanzeng, et al.
Veröffentlicht: (2025)
von: Li, Yanzeng, et al.
Veröffentlicht: (2025)
Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding
von: Shoham, Ofir Ben
Veröffentlicht: (2026)
von: Shoham, Ofir Ben
Veröffentlicht: (2026)
Judge Decoding: Faster Speculative Sampling Requires Going Beyond Model Alignment
von: Bachmann, Gregor, et al.
Veröffentlicht: (2025)
von: Bachmann, Gregor, et al.
Veröffentlicht: (2025)
Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
von: Ziashahabi, Amir, et al.
Veröffentlicht: (2025)
von: Ziashahabi, Amir, et al.
Veröffentlicht: (2025)
DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
von: Zhang, Jinbin, et al.
Veröffentlicht: (2025)
von: Zhang, Jinbin, et al.
Veröffentlicht: (2025)
SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding
von: Bang, Jehyeon, et al.
Veröffentlicht: (2026)
von: Bang, Jehyeon, et al.
Veröffentlicht: (2026)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2025)
von: Li, Jinze, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ParallelSpec: Parallel Drafter for Efficient Speculative Decoding
von: Xiao, Zilin, et al.
Veröffentlicht: (2024) -
MoE-Spec: Expert Budgeting for Efficient Speculative Decoding
von: McDanel, Bradley, et al.
Veröffentlicht: (2026) -
EasySpec: Layer-Parallel Speculative Decoding for Efficient Multi-GPU Utilization
von: Wu, Yize, et al.
Veröffentlicht: (2025) -
SpecPipe: Accelerating Pipeline Parallelism-based LLM Inference with Speculative Decoding
von: Yin, Haofei, et al.
Veröffentlicht: (2025) -
HiSpec: Hierarchical Speculative Decoding for LLMs
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)