SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Yuhao, Shen, Junyi, Kong, Quan, Liu, Tianyu, Lu, Yao, Wang, Cong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
SpecMemo: Speculative Decoding is in Your Pocket
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
MineDraft: A Framework for Batch Parallel Speculative Decoding
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
von: Pan, Rui, et al.
Veröffentlicht: (2025)
von: Pan, Rui, et al.
Veröffentlicht: (2025)
Regulating Branch Parallelism in LLM Serving
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2026)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2026)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
PipeSpec: Breaking Stage Dependencies in Hierarchical LLM Decoding
von: McDanel, Bradley, et al.
Veröffentlicht: (2025)
von: McDanel, Bradley, et al.
Veröffentlicht: (2025)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios
von: Hu, Xinyi, et al.
Veröffentlicht: (2026)
von: Hu, Xinyi, et al.
Veröffentlicht: (2026)
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
von: Shukla, Shikhar
Veröffentlicht: (2026)
von: Shukla, Shikhar
Veröffentlicht: (2026)
Faster Vertex Cover Algorithms on GPUs with Component-Aware Parallel Branching
von: Amro, Hussein, et al.
Veröffentlicht: (2025)
von: Amro, Hussein, et al.
Veröffentlicht: (2025)
Make Every Draft Count: Hidden State based Speculative Decoding
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
von: Chen, Yuetao, et al.
Veröffentlicht: (2026)
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
von: Zheng, Ce, et al.
Veröffentlicht: (2026)
von: Zheng, Ce, et al.
Veröffentlicht: (2026)
SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
Batch Query Processing and Optimization for Agentic Workflows
von: Shen, Junyi, et al.
Veröffentlicht: (2025)
von: Shen, Junyi, et al.
Veröffentlicht: (2025)
SpecInF: Exploiting Idle GPU Resources in Distributed DL Training via Speculative Inference Filling
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
von: Abramovich, Talor, et al.
Veröffentlicht: (2026)
von: Abramovich, Talor, et al.
Veröffentlicht: (2026)
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Dora: QoE-Aware Hybrid Parallelism for Distributed Edge AI
von: Jin, Jianli, et al.
Veröffentlicht: (2025)
von: Jin, Jianli, et al.
Veröffentlicht: (2025)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
Act While Thinking: Accelerating LLM Agents via Pattern-Aware Speculative Tool Execution
von: Sui, Yifan, et al.
Veröffentlicht: (2026)
von: Sui, Yifan, et al.
Veröffentlicht: (2026)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
Hybrid-Parallel: Achieving High Performance and Energy Efficient Distributed Inference on Robots
von: Sun, Zekai, et al.
Veröffentlicht: (2024)
von: Sun, Zekai, et al.
Veröffentlicht: (2024)
S-HPLB: Efficient LLM Attention Serving via Sparsity-Aware Head Parallelism Load Balance
von: Liu, Di, et al.
Veröffentlicht: (2026)
von: Liu, Di, et al.
Veröffentlicht: (2026)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
Tasking framework for Adaptive Speculative Parallel Mesh Generation
von: Tsolakis, Christos, et al.
Veröffentlicht: (2024)
von: Tsolakis, Christos, et al.
Veröffentlicht: (2024)
Distributed Semi-Speculative Parallel Anisotropic Mesh Adaptation
von: Garner, Kevin, et al.
Veröffentlicht: (2026)
von: Garner, Kevin, et al.
Veröffentlicht: (2026)
TEE is not a Healer: Rollback-Resistant Reliable Storage (Extended Version)
von: Keshavarzi, Sadegh, et al.
Veröffentlicht: (2025)
von: Keshavarzi, Sadegh, et al.
Veröffentlicht: (2025)
Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
Utility-Driven Speculative Decoding for Mixture-of-Experts
von: Saxena, Anish, et al.
Veröffentlicht: (2025)
von: Saxena, Anish, et al.
Veröffentlicht: (2025)
Distributed Speculative Execution for Resilient Cloud Applications
von: Li, Tianyu, et al.
Veröffentlicht: (2024)
von: Li, Tianyu, et al.
Veröffentlicht: (2024)
Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
von: Cheng, Rongxin, et al.
Veröffentlicht: (2025)
von: Cheng, Rongxin, et al.
Veröffentlicht: (2025)
SimpleFSDP: Simpler Fully Sharded Data Parallel with torch.compile
von: Zhang, Ruisi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruisi, et al.
Veröffentlicht: (2024)
B-PASTE: Beam-Aware Pattern-Guided Speculative Execution for Resource-Constrained LLM Agents
von: Song, Yanfei
Veröffentlicht: (2026)
von: Song, Yanfei
Veröffentlicht: (2026)
Ähnliche Einträge
-
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
von: Li, Yuchen, et al.
Veröffentlicht: (2026) -
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025) -
SpecMemo: Speculative Decoding is in Your Pocket
von: Yildirim, Selin, et al.
Veröffentlicht: (2025) -
MineDraft: A Framework for Batch Parallel Speculative Decoding
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026) -
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)