MineDraft: A Framework for Batch Parallel Speculative Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Zhenwei, Verma, Arun, Zhou, Zijian, Wu, Zhaoxuan, Prakash, Alok, Rus, Daniela, Low, Bryan Kian Hsiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
Tasking framework for Adaptive Speculative Parallel Mesh Generation
von: Tsolakis, Christos, et al.
Veröffentlicht: (2024)
von: Tsolakis, Christos, et al.
Veröffentlicht: (2024)
Distributed Semi-Speculative Parallel Anisotropic Mesh Adaptation
von: Garner, Kevin, et al.
Veröffentlicht: (2026)
von: Garner, Kevin, et al.
Veröffentlicht: (2026)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
Herring: Parallel Batch-Order-Fairness on DAG-based Blockchain Consensus
von: Putnik, Marko, et al.
Veröffentlicht: (2026)
von: Putnik, Marko, et al.
Veröffentlicht: (2026)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
Make Every Draft Count: Hidden State based Speculative Decoding
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
von: Sakip, Akhmed, et al.
Veröffentlicht: (2026)
von: Sakip, Akhmed, et al.
Veröffentlicht: (2026)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding
von: Chen, Jiefei, et al.
Veröffentlicht: (2026)
von: Chen, Jiefei, et al.
Veröffentlicht: (2026)
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
von: Pan, Rui, et al.
Veröffentlicht: (2025)
von: Pan, Rui, et al.
Veröffentlicht: (2025)
Parallelizing Maximal Clique Enumeration on GPUs
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2026)
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2026)
Parallel Online Directed Acyclic Graph Exploration for Atlasing Soft-Matter Assembly Configuration Spaces
von: Prabhu, Rahul, et al.
Veröffentlicht: (2024)
von: Prabhu, Rahul, et al.
Veröffentlicht: (2024)
Parallel Data Object Creation: Towards Scalable Metadata Management in High-Performance I/O Library
von: Li, Youjia, et al.
Veröffentlicht: (2025)
von: Li, Youjia, et al.
Veröffentlicht: (2025)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
von: Tang, Ding, et al.
Veröffentlicht: (2024)
von: Tang, Ding, et al.
Veröffentlicht: (2024)
Distributed Speculative Execution for Resilient Cloud Applications
von: Li, Tianyu, et al.
Veröffentlicht: (2024)
von: Li, Tianyu, et al.
Veröffentlicht: (2024)
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
von: Abramovich, Talor, et al.
Veröffentlicht: (2026)
von: Abramovich, Talor, et al.
Veröffentlicht: (2026)
GPU-Accelerated Batch-Dynamic Subgraph Matching
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control
von: Chen, Qiaoling, et al.
Veröffentlicht: (2026)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2026)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Balancing Pipeline Parallelism with Vocabulary Parallelism
von: Yeung, Man Tsung, et al.
Veröffentlicht: (2024)
von: Yeung, Man Tsung, et al.
Veröffentlicht: (2024)
Revisiting Speculative Leaderless Protocols for Low-Latency BFT Replication
von: Qian, Daniel, et al.
Veröffentlicht: (2026)
von: Qian, Daniel, et al.
Veröffentlicht: (2026)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
Parallel Batch-Dynamic Maximal Independent Set
von: Blelloch, Guy, et al.
Veröffentlicht: (2026)
von: Blelloch, Guy, et al.
Veröffentlicht: (2026)
Scalable and Cost-Efficient ML Inference: Parallel Batch Processing with Serverless Functions
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
von: Barrak, Amine, et al.
Veröffentlicht: (2025)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
von: Xu, Yaodan, et al.
Veröffentlicht: (2025)
von: Xu, Yaodan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
von: Wu, Zhaoxuan, et al.
Veröffentlicht: (2025) -
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
von: Shen, Yuhao, et al.
Veröffentlicht: (2025) -
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
von: Li, Yuchen, et al.
Veröffentlicht: (2026) -
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
von: Wei, Jinhui, et al.
Veröffentlicht: (2025) -
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)