Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Timor, Nadav, Mamou, Jonathan, Korat, Daniel, Berchansky, Moshe, Jain, Gaurav, Pereg, Oren, Wasserblat, Moshe, Harel, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dynamic Speculation Lookahead Accelerates Speculative Decoding of Large Language Models
von: Mamou, Jonathan, et al.
Veröffentlicht: (2024)
von: Mamou, Jonathan, et al.
Veröffentlicht: (2024)
Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference
von: Timor, Nadav, et al.
Veröffentlicht: (2024)
von: Timor, Nadav, et al.
Veröffentlicht: (2024)
Out-of-Vocabulary Sampling Boosts Speculative Decoding
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
CoTAR: Chain-of-Thought Attribution Reasoning with Multi-level Granularity
von: Berchansky, Moshe, et al.
Veröffentlicht: (2024)
von: Berchansky, Moshe, et al.
Veröffentlicht: (2024)
SQuARE: Sequential Question Answering Reasoning Engine for Enhanced Chain-of-Thought in Large Language Models
von: Fleischer, Daniel, et al.
Veröffentlicht: (2025)
von: Fleischer, Daniel, et al.
Veröffentlicht: (2025)
RAG Foundry: A Framework for Enhancing LLMs for Retrieval Augmented Generation
von: Fleischer, Daniel, et al.
Veröffentlicht: (2024)
von: Fleischer, Daniel, et al.
Veröffentlicht: (2024)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
Speculative Decoding with a Speculative Vocabulary
von: Williams, Miles, et al.
Veröffentlicht: (2026)
von: Williams, Miles, et al.
Veröffentlicht: (2026)
SWIFT: On-the-Fly Self-Speculative Decoding for LLM Inference Acceleration
von: Xia, Heming, et al.
Veröffentlicht: (2024)
von: Xia, Heming, et al.
Veröffentlicht: (2024)
Speculate Deep and Accurate: Lossless and Training-Free Acceleration for Offloaded LLMs via Substitute Speculative Decoding
von: Wang, Pei-Shuo, et al.
Veröffentlicht: (2025)
von: Wang, Pei-Shuo, et al.
Veröffentlicht: (2025)
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
von: Sun, Hanshi, et al.
Veröffentlicht: (2024)
Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
von: Zhang, Jun, et al.
Veröffentlicht: (2023)
von: Zhang, Jun, et al.
Veröffentlicht: (2023)
NanoSpec: Accelerating Speculative Decoding using Minimalist In-Context Vocabularies
von: Chen, Zhiyang, et al.
Veröffentlicht: (2026)
von: Chen, Zhiyang, et al.
Veröffentlicht: (2026)
Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding
von: Cho, Sukmin, et al.
Veröffentlicht: (2025)
von: Cho, Sukmin, et al.
Veröffentlicht: (2025)
SDSAT: Accelerating LLM Inference through Speculative Decoding with Semantic Adaptive Tokens
von: Liu, Chengbo, et al.
Veröffentlicht: (2024)
von: Liu, Chengbo, et al.
Veröffentlicht: (2024)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
Tutorial Proposal: Speculative Decoding for Efficient LLM Inference
von: Xia, Heming, et al.
Veröffentlicht: (2025)
von: Xia, Heming, et al.
Veröffentlicht: (2025)
Automatic Task Detection and Heterogeneous LLM Speculative Decoding
von: Ge, Danying, et al.
Veröffentlicht: (2025)
von: Ge, Danying, et al.
Veröffentlicht: (2025)
Kangaroo: Lossless Self-Speculative Decoding via Double Early Exiting
von: Liu, Fangcheng, et al.
Veröffentlicht: (2024)
von: Liu, Fangcheng, et al.
Veröffentlicht: (2024)
Hierarchical Verification of Speculative Beams for Accelerating LLM Inference
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
DART: Diffusion-Inspired Speculative Decoding for Fast LLM Inference
von: Liu, Fuliang, et al.
Veröffentlicht: (2026)
von: Liu, Fuliang, et al.
Veröffentlicht: (2026)
HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
von: Yen, Howard, et al.
Veröffentlicht: (2024)
von: Yen, Howard, et al.
Veröffentlicht: (2024)
Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inference
von: Zhang, Libo, et al.
Veröffentlicht: (2024)
von: Zhang, Libo, et al.
Veröffentlicht: (2024)
VOCABTRIM: Vocabulary Pruning for Efficient Speculative Decoding in LLMs
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
von: Goel, Raghavv, et al.
Veröffentlicht: (2025)
Boosting Lossless Speculative Decoding via Feature Sampling and Partial Alignment Distillation
von: Gui, Lujun, et al.
Veröffentlicht: (2024)
von: Gui, Lujun, et al.
Veröffentlicht: (2024)
Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters
von: Yi, Euiin, et al.
Veröffentlicht: (2024)
von: Yi, Euiin, et al.
Veröffentlicht: (2024)
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
von: Yang, Penghui, et al.
Veröffentlicht: (2025)
Adapting LLMs to Hebrew: Unveiling DictaLM 2.0 with Enhanced Vocabulary and Instruction Capabilities
von: Shmidman, Shaltiel, et al.
Veröffentlicht: (2024)
von: Shmidman, Shaltiel, et al.
Veröffentlicht: (2024)
Accelerate Speculative Decoding with Sparse Computation in Verification
von: Wang, Jikai, et al.
Veröffentlicht: (2025)
von: Wang, Jikai, et al.
Veröffentlicht: (2025)
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
von: McDanel, Bradley
Veröffentlicht: (2024)
von: McDanel, Bradley
Veröffentlicht: (2024)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
Chimera: A Lossless Decoding Method for Accelerating Large Language Models Inference by Fusing all Tokens
von: Zeng, Ziqian, et al.
Veröffentlicht: (2024)
von: Zeng, Ziqian, et al.
Veröffentlicht: (2024)
SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
von: Svirschevski, Ruslan, et al.
Veröffentlicht: (2024)
von: Svirschevski, Ruslan, et al.
Veröffentlicht: (2024)
Accelerating Speculative Decoding with Block Diffusion Draft Trees
von: Ringel, Liran, et al.
Veröffentlicht: (2026)
von: Ringel, Liran, et al.
Veröffentlicht: (2026)
Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decoding
von: Shoham, Ofir Ben
Veröffentlicht: (2026)
von: Shoham, Ofir Ben
Veröffentlicht: (2026)
TALL -- A Trainable Architecture for Enhancing LLM Performance in Low-Resource Languages
von: Ofer, Moshe, et al.
Veröffentlicht: (2025)
von: Ofer, Moshe, et al.
Veröffentlicht: (2025)
HeteroSpec: Leveraging Contextual Heterogeneity for Efficient Speculative Decoding
von: Liu, Siran, et al.
Veröffentlicht: (2025)
von: Liu, Siran, et al.
Veröffentlicht: (2025)
Block Verification Accelerates Speculative Decoding
von: Sun, Ziteng, et al.
Veröffentlicht: (2024)
von: Sun, Ziteng, et al.
Veröffentlicht: (2024)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dynamic Speculation Lookahead Accelerates Speculative Decoding of Large Language Models
von: Mamou, Jonathan, et al.
Veröffentlicht: (2024) -
Distributed Speculative Inference (DSI): Speculation Parallelism for Provably Faster Lossless Language Model Inference
von: Timor, Nadav, et al.
Veröffentlicht: (2024) -
Out-of-Vocabulary Sampling Boosts Speculative Decoding
von: Timor, Nadav, et al.
Veröffentlicht: (2025) -
CoTAR: Chain-of-Thought Attribution Reasoning with Multi-level Granularity
von: Berchansky, Moshe, et al.
Veröffentlicht: (2024) -
SQuARE: Sequential Question Answering Reasoning Engine for Enhanced Chain-of-Thought in Large Language Models
von: Fleischer, Daniel, et al.
Veröffentlicht: (2025)