Multi-Draft Speculative Sampling: Canonical Decomposition and Theoretical Limits
Fuente:
arXiv
Saved in:
| Main Authors: | Khisti, Ashish, Ebrahimi, M. Reza, Dbouk, Hassan, Behboodi, Arash, Memisevic, Roland, Louizos, Christos |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SECO: Secure Inference With Model Splitting Across Multi-Server Hierarchy
by: Chen, Shuangyi, et al.
Published: (2024)
by: Chen, Shuangyi, et al.
Published: (2024)
Tasking framework for Adaptive Speculative Parallel Mesh Generation
by: Tsolakis, Christos, et al.
Published: (2024)
by: Tsolakis, Christos, et al.
Published: (2024)
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
by: Li, Yuchen, et al.
Published: (2026)
by: Li, Yuchen, et al.
Published: (2026)
Robust Federated Finetuning of Foundation Models via Alternating Minimization of LoRA
by: Chen, Shuangyi, et al.
Published: (2024)
by: Chen, Shuangyi, et al.
Published: (2024)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
by: Shen, Yuhao, et al.
Published: (2025)
by: Shen, Yuhao, et al.
Published: (2025)
FedMIL: Federated-Multiple Instance Learning for Video Analysis with Optimized DPP Scheduling
by: Bastola, Ashish, et al.
Published: (2024)
by: Bastola, Ashish, et al.
Published: (2024)
Fundamental Limits of Coded Polynomial Aggregation
by: Zhong, Xi, et al.
Published: (2026)
by: Zhong, Xi, et al.
Published: (2026)
Stable Diffusion-based Data Augmentation for Federated Learning with Non-IID Data
by: Morafah, Mahdi, et al.
Published: (2024)
by: Morafah, Mahdi, et al.
Published: (2024)
Distributed Speculative Execution for Resilient Cloud Applications
by: Li, Tianyu, et al.
Published: (2024)
by: Li, Tianyu, et al.
Published: (2024)
Distributed Semi-Speculative Parallel Anisotropic Mesh Adaptation
by: Garner, Kevin, et al.
Published: (2026)
by: Garner, Kevin, et al.
Published: (2026)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
by: Dai, Yuntao, et al.
Published: (2026)
by: Dai, Yuntao, et al.
Published: (2026)
Revisiting Speculative Leaderless Protocols for Low-Latency BFT Replication
by: Qian, Daniel, et al.
Published: (2026)
by: Qian, Daniel, et al.
Published: (2026)
Perfect Multi-User Distributed Computing
by: Khalesi, Ali, et al.
Published: (2024)
by: Khalesi, Ali, et al.
Published: (2024)
GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations
by: Yousefzadeh-Asl-Miandoab, Ehsan, et al.
Published: (2026)
by: Yousefzadeh-Asl-Miandoab, Ehsan, et al.
Published: (2026)
MineDraft: A Framework for Batch Parallel Speculative Decoding
by: Tang, Zhenwei, et al.
Published: (2026)
by: Tang, Zhenwei, et al.
Published: (2026)
Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution
by: Wang, Siqi, et al.
Published: (2024)
by: Wang, Siqi, et al.
Published: (2024)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
by: Wang, Zhibin, et al.
Published: (2025)
by: Wang, Zhibin, et al.
Published: (2025)
SPIN: Accelerating Large Language Model Inference with Heterogeneous Speculative Models
by: Chen, Fahao, et al.
Published: (2025)
by: Chen, Fahao, et al.
Published: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
by: Wei, Jinhui, et al.
Published: (2025)
by: Wei, Jinhui, et al.
Published: (2025)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
by: Chen, Wenyan, et al.
Published: (2026)
by: Chen, Wenyan, et al.
Published: (2026)
Path Connected Dynamic Graphs with a Study of Dispersion and Exploration
by: Saxena, Ashish, et al.
Published: (2025)
by: Saxena, Ashish, et al.
Published: (2025)
Exploration on Highly Dynamic Graphs
by: Saxena, Ashish, et al.
Published: (2026)
by: Saxena, Ashish, et al.
Published: (2026)
When Agents are Powerful: Black Hole Search with Verification in Time-Varying Graphs
by: Kaur, Tanvir, et al.
Published: (2025)
by: Kaur, Tanvir, et al.
Published: (2025)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
by: Zhang, Yida, et al.
Published: (2026)
by: Zhang, Yida, et al.
Published: (2026)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
by: Tran, Phuong, et al.
Published: (2025)
by: Tran, Phuong, et al.
Published: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
by: Han, Yunhe, et al.
Published: (2026)
by: Han, Yunhe, et al.
Published: (2026)
Sparse Checkpointing for Fast and Reliable MoE Training
by: Gandhi, Swapnil, et al.
Published: (2024)
by: Gandhi, Swapnil, et al.
Published: (2024)
Make Every Draft Count: Hidden State based Speculative Decoding
by: Chen, Yuetao, et al.
Published: (2026)
by: Chen, Yuetao, et al.
Published: (2026)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
by: Chen, Liangkun, et al.
Published: (2025)
by: Chen, Liangkun, et al.
Published: (2025)
SpecInF: Exploiting Idle GPU Resources in Distributed DL Training via Speculative Inference Filling
by: Lv, Cunchi, et al.
Published: (2025)
by: Lv, Cunchi, et al.
Published: (2025)
MIRAGE: Runtime Scheduling for Multi-Vector Image Retrieval with Hierarchical Decomposition
by: Li, Maoliang, et al.
Published: (2025)
by: Li, Maoliang, et al.
Published: (2025)
Balanced Dispersion on Time-Varying Dynamic Graphs
by: Saxena, Ashish, et al.
Published: (2024)
by: Saxena, Ashish, et al.
Published: (2024)
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
Experimental Evaluation of Distributed k-Core Decomposition
by: Guo, Bin, et al.
Published: (2024)
by: Guo, Bin, et al.
Published: (2024)
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA Graphs
by: Sojoodi, Amirhossein, et al.
Published: (2026)
by: Sojoodi, Amirhossein, et al.
Published: (2026)
Performance and Security Aware Distributed Service Placement in Fog Computing
by: Goudarzi, Mohammad, et al.
Published: (2026)
by: Goudarzi, Mohammad, et al.
Published: (2026)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
Logarithmic-Time Geodesically Convex Decomposition in Programmable Matter
by: Hillebrandt, Henning, et al.
Published: (2026)
by: Hillebrandt, Henning, et al.
Published: (2026)
Federated Learning Using Coupled Tensor Train Decomposition
by: Zhang, Xiangtao, et al.
Published: (2024)
by: Zhang, Xiangtao, et al.
Published: (2024)
PRISM: Processing-In-Memory Sparse MTTKRP for Tensor Decomposition Acceleration
by: Pacheco, Daniel, et al.
Published: (2026)
by: Pacheco, Daniel, et al.
Published: (2026)
Similar Items
-
SECO: Secure Inference With Model Splitting Across Multi-Server Hierarchy
by: Chen, Shuangyi, et al.
Published: (2024) -
Tasking framework for Adaptive Speculative Parallel Mesh Generation
by: Tsolakis, Christos, et al.
Published: (2024) -
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
by: Li, Yuchen, et al.
Published: (2026) -
Robust Federated Finetuning of Foundation Models via Alternating Minimization of LoRA
by: Chen, Shuangyi, et al.
Published: (2024) -
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
by: Shen, Yuhao, et al.
Published: (2025)