Overcoming Joint Intractability with Lossless Hierarchical Speculative Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Yuxuan, Huang, Fei, Li, Heng, Wu, Fengyi, Wang, Tianyu, Zhang, Jianwei, Lin, Junyang, Cheng, Zhi-Qi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
von: Ning, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Ning, Zhiyuan, et al.
Veröffentlicht: (2025)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
MaxSup: Overcoming Representation Collapse in Label Smoothing
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2025)
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2025)
SAM Decoding: Speculative Decoding via Suffix Automaton
von: Hu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2024)
OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs
von: Lee, Jaeseong, et al.
Veröffentlicht: (2025)
von: Lee, Jaeseong, et al.
Veröffentlicht: (2025)
Batch Speculative Decoding Done Right
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
HiSpec: Hierarchical Speculative Decoding for LLMs
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
von: Kumar, Avinash, et al.
Veröffentlicht: (2025)
Speculative Safety-Aware Decoding
von: Wang, Xuekang, et al.
Veröffentlicht: (2025)
von: Wang, Xuekang, et al.
Veröffentlicht: (2025)
HIPPO: Accelerating Video Large Language Models Inference via Holistic-aware Parallel Speculative Decoding
von: Lv, Qitan, et al.
Veröffentlicht: (2026)
von: Lv, Qitan, et al.
Veröffentlicht: (2026)
HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions
von: Dong, Yifei, et al.
Veröffentlicht: (2025)
von: Dong, Yifei, et al.
Veröffentlicht: (2025)
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
von: Huang, Zongle, et al.
Veröffentlicht: (2025)
von: Huang, Zongle, et al.
Veröffentlicht: (2025)
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
Cacheback: Speculative Decoding With Nothing But Cache
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms
von: Hou, Yunlong, et al.
Veröffentlicht: (2025)
von: Hou, Yunlong, et al.
Veröffentlicht: (2025)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2025)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design
von: Zhang, Yudi, et al.
Veröffentlicht: (2025)
von: Zhang, Yudi, et al.
Veröffentlicht: (2025)
Free Draft-and-Verification: Toward Lossless Parallel Decoding for Diffusion Large Language Models
von: Wu, Shutong, et al.
Veröffentlicht: (2025)
von: Wu, Shutong, et al.
Veröffentlicht: (2025)
Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding
von: Shen, Yuhao, et al.
Veröffentlicht: (2026)
von: Shen, Yuhao, et al.
Veröffentlicht: (2026)
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2025)
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
Lossless Model Compression via Joint Low-Rank Factorization Optimization
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
von: Zhang, Boyang, et al.
Veröffentlicht: (2024)
Fuzzy Speculative Decoding for a Tunable Accuracy-Runtime Tradeoff
von: Holsman, Maximilian, et al.
Veröffentlicht: (2025)
von: Holsman, Maximilian, et al.
Veröffentlicht: (2025)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
GRIFFIN: Effective Token Alignment for Faster Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
RASD: Retrieval-Augmented Speculative Decoding
von: Quan, Guofeng, et al.
Veröffentlicht: (2025)
von: Quan, Guofeng, et al.
Veröffentlicht: (2025)
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
QSpec: Speculative Decoding with Complementary Quantization Schemes
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
The Disparate Impacts of Speculative Decoding
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
Constrained Decoding with Speculative Lookaheads
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
Cross-Attention Speculative Decoding
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
Speculative Decoding: Performance or Illusion?
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
Mamba Drafters for Speculative Decoding
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
Speculative Actions: A Lossless Framework for Faster Agentic Systems
von: Ye, Naimeng, et al.
Veröffentlicht: (2025)
von: Ye, Naimeng, et al.
Veröffentlicht: (2025)
Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding
von: Jin, Tao, et al.
Veröffentlicht: (2026)
von: Jin, Tao, et al.
Veröffentlicht: (2026)
Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
von: Hu, Shijing, et al.
Veröffentlicht: (2025)
Reinforcement Speculative Decoding for Fast Ranking
von: Du, Yingpeng, et al.
Veröffentlicht: (2025)
von: Du, Yingpeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
von: Yang, Penghui, et al.
Veröffentlicht: (2025) -
CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMs
von: Ning, Zhiyuan, et al.
Veröffentlicht: (2025) -
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025) -
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
von: Timor, Nadav, et al.
Veröffentlicht: (2025) -
MaxSup: Overcoming Representation Collapse in Label Smoothing
von: Zhou, Yuxuan, et al.
Veröffentlicht: (2025)