Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
Fuente:
arXiv
Guardado en:
| Autores principales: | Metel, Michael R., Lu, Peng, Chen, Boxing, Rezagholizadeh, Mehdi, Kobyzev, Ivan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
por: Metel, Michael R., et al.
Publicado: (2024)
por: Metel, Michael R., et al.
Publicado: (2024)
ReGLA: Refining Gated Linear Attention
por: Lu, Peng, et al.
Publicado: (2025)
por: Lu, Peng, et al.
Publicado: (2025)
OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
por: Huang, Chenyang, et al.
Publicado: (2024)
por: Huang, Chenyang, et al.
Publicado: (2024)
Resonance RoPE: Improving Context Length Generalization of Large Language Models
por: Wang, Suyuchen, et al.
Publicado: (2024)
por: Wang, Suyuchen, et al.
Publicado: (2024)
LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization
por: Lu, Peng, et al.
Publicado: (2023)
por: Lu, Peng, et al.
Publicado: (2023)
S2D: Sorted Speculative Decoding For More Efficient Deployment of Nested Large Language Models
por: Kavehzadeh, Parsa, et al.
Publicado: (2024)
por: Kavehzadeh, Parsa, et al.
Publicado: (2024)
Integral Transformer: Denoising Attention, Not Too Much Not Too Little
por: Kobyzev, Ivan, et al.
Publicado: (2025)
por: Kobyzev, Ivan, et al.
Publicado: (2025)
BOSCH: Black-Box Binary Optimization for Short-Context Attention-Head Selection in LLMs
por: Ghaddar, Abbas, et al.
Publicado: (2026)
por: Ghaddar, Abbas, et al.
Publicado: (2026)
On the importance of Data Scale in Pretraining Arabic Language Models
por: Ghaddar, Abbas, et al.
Publicado: (2024)
por: Ghaddar, Abbas, et al.
Publicado: (2024)
PEARL: Parallel Speculative Decoding with Adaptive Draft Length
por: Liu, Tianyu, et al.
Publicado: (2024)
por: Liu, Tianyu, et al.
Publicado: (2024)
Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning
por: Zhang, Jiebin, et al.
Publicado: (2026)
por: Zhang, Jiebin, et al.
Publicado: (2026)
Thinking Long, but Short: Stable Sequential Test-Time Scaling for Large Reasoning Models
por: Metel, Michael R., et al.
Publicado: (2026)
por: Metel, Michael R., et al.
Publicado: (2026)
OPT-Tree: Speculative Decoding with Adaptive Draft Tree Structure
por: Wang, Jikai, et al.
Publicado: (2024)
por: Wang, Jikai, et al.
Publicado: (2024)
AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures
por: Zhang, Situo, et al.
Publicado: (2024)
por: Zhang, Situo, et al.
Publicado: (2024)
Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
por: Zhang, Jun, et al.
Publicado: (2023)
por: Zhang, Jun, et al.
Publicado: (2023)
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration
por: Xia, Heming, et al.
Publicado: (2024)
por: Xia, Heming, et al.
Publicado: (2024)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
por: Huang, Jerry, et al.
Publicado: (2024)
por: Huang, Jerry, et al.
Publicado: (2024)
Cost-Aware Diffusion Draft Trees for Speculative Decoding
por: Zhang, Shuai, et al.
Publicado: (2026)
por: Zhang, Shuai, et al.
Publicado: (2026)
Accelerating Speculative Decoding with Block Diffusion Draft Trees
por: Ringel, Liran, et al.
Publicado: (2026)
por: Ringel, Liran, et al.
Publicado: (2026)
OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding
por: Ramakrishnan, Ramchalam Kinattinkara, et al.
Publicado: (2025)
por: Ramakrishnan, Ramchalam Kinattinkara, et al.
Publicado: (2025)
TABED: Test-Time Adaptive Ensemble Drafting for Robust Speculative Decoding in LVLMs
por: Lee, Minjae, et al.
Publicado: (2026)
por: Lee, Minjae, et al.
Publicado: (2026)
Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding
por: Hu, Shijing, et al.
Publicado: (2025)
por: Hu, Shijing, et al.
Publicado: (2025)
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding
por: Sun, Ryan, et al.
Publicado: (2024)
por: Sun, Ryan, et al.
Publicado: (2024)
CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems
por: Ghaddar, Abbas, et al.
Publicado: (2024)
por: Ghaddar, Abbas, et al.
Publicado: (2024)
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
por: Kavehzadeh, Parsa, et al.
Publicado: (2023)
por: Kavehzadeh, Parsa, et al.
Publicado: (2023)
SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting
por: Shi, Weijie, et al.
Publicado: (2026)
por: Shi, Weijie, et al.
Publicado: (2026)
Make Every Draft Count: Hidden State based Speculative Decoding
por: Chen, Yuetao, et al.
Publicado: (2026)
por: Chen, Yuetao, et al.
Publicado: (2026)
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
por: Zhang, Ziyin, et al.
Publicado: (2024)
por: Zhang, Ziyin, et al.
Publicado: (2024)
DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting
por: Lv, Kai, et al.
Publicado: (2025)
por: Lv, Kai, et al.
Publicado: (2025)
DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding
por: Hu, Yunhai, et al.
Publicado: (2025)
por: Hu, Yunhai, et al.
Publicado: (2025)
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
por: Zhao, Weilin, et al.
Publicado: (2024)
por: Zhao, Weilin, et al.
Publicado: (2024)
Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
por: Huang, Jianuo, et al.
Publicado: (2026)
por: Huang, Jianuo, et al.
Publicado: (2026)
DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding
por: Zhang, Jiebin, et al.
Publicado: (2026)
por: Zhang, Jiebin, et al.
Publicado: (2026)
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
por: Wu, Zhaoxuan, et al.
Publicado: (2025)
por: Wu, Zhaoxuan, et al.
Publicado: (2025)
Draft-OPD: On-Policy Distillation for Speculative Draft Models
por: Lei, Haodi, et al.
Publicado: (2026)
por: Lei, Haodi, et al.
Publicado: (2026)
SpecTr-GBV: Multi-Draft Block Verification Accelerating Speculative Decoding
por: Lin, Yijun, et al.
Publicado: (2026)
por: Lin, Yijun, et al.
Publicado: (2026)
KOALA: Enhancing Speculative Decoding for LLM via Multi-Layer Draft Heads with Adversarial Learning
por: Zhang, Kaiqi, et al.
Publicado: (2024)
por: Zhang, Kaiqi, et al.
Publicado: (2024)
POSS: Position Specialist Generates Better Draft for Speculative Decoding
por: Huang, Langlin, et al.
Publicado: (2025)
por: Huang, Langlin, et al.
Publicado: (2025)
MineDraft: A Framework for Batch Parallel Speculative Decoding
por: Tang, Zhenwei, et al.
Publicado: (2026)
por: Tang, Zhenwei, et al.
Publicado: (2026)
FlexDraft: Flexible Speculative Decoding via Attention Tuning and Bonus-Guided Calibration
por: Zhang, Yaojie, et al.
Publicado: (2026)
por: Zhang, Yaojie, et al.
Publicado: (2026)
Ejemplares similares
-
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
por: Metel, Michael R., et al.
Publicado: (2024) -
ReGLA: Refining Gated Linear Attention
por: Lu, Peng, et al.
Publicado: (2025) -
OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
por: Huang, Chenyang, et al.
Publicado: (2024) -
Resonance RoPE: Improving Context Length Generalization of Large Language Models
por: Wang, Suyuchen, et al.
Publicado: (2024) -
LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization
por: Lu, Peng, et al.
Publicado: (2023)