3-Model Speculative Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Byun, Sanghyun, Odema, Mohanad, Guack, Jung Ick, Lee, Baisub, Song, Jacob, Chung, Woo Seong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
APCE: Adaptive Progressive Context Expansion for Long Context Processing
von: Lee, Baisub, et al.
Veröffentlicht: (2025)
von: Lee, Baisub, et al.
Veröffentlicht: (2025)
MultiDepth: Multi-Sample Priors for Refining Monocular Metric Depth Estimations in Indoor Scenes
von: Byun, Sanghyun, et al.
Veröffentlicht: (2024)
von: Byun, Sanghyun, et al.
Veröffentlicht: (2024)
Training-free Dropout Sampling for Semantic Token Acceptance in Speculative Decoding
von: Lee, Jeongtae, et al.
Veröffentlicht: (2026)
von: Lee, Jeongtae, et al.
Veröffentlicht: (2026)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)
Multi-Drafter Speculative Decoding with Alignment Feedback
von: Kim, Taehyeon, et al.
Veröffentlicht: (2026)
von: Kim, Taehyeon, et al.
Veröffentlicht: (2026)
OneNet: A Channel-Wise 1D Convolutional U-Net
von: Byun, Sanghyun, et al.
Veröffentlicht: (2024)
von: Byun, Sanghyun, et al.
Veröffentlicht: (2024)
Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding
von: Kim, Sungkyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungkyun, et al.
Veröffentlicht: (2025)
Speculative Decoding with a Speculative Vocabulary
von: Williams, Miles, et al.
Veröffentlicht: (2026)
von: Williams, Miles, et al.
Veröffentlicht: (2026)
Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters
von: Yi, Euiin, et al.
Veröffentlicht: (2024)
von: Yi, Euiin, et al.
Veröffentlicht: (2024)
Decoding Speculative Decoding
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
von: Yan, Minghao, et al.
Veröffentlicht: (2024)
Dynamic Speculation Lookahead Accelerates Speculative Decoding of Large Language Models
von: Mamou, Jonathan, et al.
Veröffentlicht: (2024)
von: Mamou, Jonathan, et al.
Veröffentlicht: (2024)
Cross-Attention Speculative Decoding
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
Speculative Contrastive Decoding
von: Yuan, Hongyi, et al.
Veröffentlicht: (2023)
von: Yuan, Hongyi, et al.
Veröffentlicht: (2023)
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
Mamba Drafters for Speculative Decoding
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
On Speculative Decoding for Multimodal Large Language Models
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
von: Gagrani, Mukul, et al.
Veröffentlicht: (2024)
SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
Multi-Candidate Speculative Decoding
von: Yang, Sen, et al.
Veröffentlicht: (2024)
von: Yang, Sen, et al.
Veröffentlicht: (2024)
Graph-Structured Speculative Decoding
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2024)
von: Gong, Zhuocheng, et al.
Veröffentlicht: (2024)
Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding
von: Cho, Sukmin, et al.
Veröffentlicht: (2025)
von: Cho, Sukmin, et al.
Veröffentlicht: (2025)
Self Speculative Decoding for Diffusion Large Language Models
von: Gao, Yifeng, et al.
Veröffentlicht: (2025)
von: Gao, Yifeng, et al.
Veröffentlicht: (2025)
A Multi-Model Adaptation of Speculative Decoding for Classification
von: Roy, Somnath, et al.
Veröffentlicht: (2025)
von: Roy, Somnath, et al.
Veröffentlicht: (2025)
Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learning
von: Zhang, Jiebin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiebin, et al.
Veröffentlicht: (2026)
Improving Multi-candidate Speculative Decoding
von: Lu, Xiaofan, et al.
Veröffentlicht: (2024)
von: Lu, Xiaofan, et al.
Veröffentlicht: (2024)
Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion
von: Christopher, Jacob K, et al.
Veröffentlicht: (2024)
von: Christopher, Jacob K, et al.
Veröffentlicht: (2024)
DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding
von: Zhang, Jiebin, et al.
Veröffentlicht: (2026)
von: Zhang, Jiebin, et al.
Veröffentlicht: (2026)
The Disparate Impacts of Speculative Decoding
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
Speculative Decoding: Performance or Illusion?
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
Constrained Decoding with Speculative Lookaheads
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
Speculative Decoding Across Languages
von: Paudel, Nirajan, et al.
Veröffentlicht: (2026)
von: Paudel, Nirajan, et al.
Veröffentlicht: (2026)
Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions
von: Tang, Bangsheng, et al.
Veröffentlicht: (2025)
von: Tang, Bangsheng, et al.
Veröffentlicht: (2025)
Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding
von: Jin, Tao, et al.
Veröffentlicht: (2026)
von: Jin, Tao, et al.
Veröffentlicht: (2026)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
Speculative Decoding and Beyond: An In-Depth Survey of Techniques
von: Hu, Yunhai, et al.
Veröffentlicht: (2025)
von: Hu, Yunhai, et al.
Veröffentlicht: (2025)
DReSD: Dense Retrieval for Speculative Decoding
von: Gritta, Milan, et al.
Veröffentlicht: (2025)
von: Gritta, Milan, et al.
Veröffentlicht: (2025)
Accelerate Speculative Decoding with Sparse Computation in Verification
von: Wang, Jikai, et al.
Veröffentlicht: (2025)
von: Wang, Jikai, et al.
Veröffentlicht: (2025)
SPEED: Speculative Pipelined Execution for Efficient Decoding
von: Hooper, Coleman, et al.
Veröffentlicht: (2023)
von: Hooper, Coleman, et al.
Veröffentlicht: (2023)
DFlash: Block Diffusion for Flash Speculative Decoding
von: Chen, Jian, et al.
Veröffentlicht: (2026)
von: Chen, Jian, et al.
Veröffentlicht: (2026)
Speculative Pipeline Decoding: Higher-Accruacy and Zero-Bubble Speculation via Pipeline Parallelism
von: Yu, Yijiong, et al.
Veröffentlicht: (2026)
von: Yu, Yijiong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025) -
APCE: Adaptive Progressive Context Expansion for Long Context Processing
von: Lee, Baisub, et al.
Veröffentlicht: (2025) -
MultiDepth: Multi-Sample Priors for Refining Monocular Metric Depth Estimations in Indoor Scenes
von: Byun, Sanghyun, et al.
Veröffentlicht: (2024) -
Training-free Dropout Sampling for Semantic Token Acceptance in Speculative Decoding
von: Lee, Jeongtae, et al.
Veröffentlicht: (2026) -
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
von: Yoon, Kanghoon, et al.
Veröffentlicht: (2025)