FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xing, Luo, Lizhuo, Tang, Ming, Huang, Chao, Chen, Xu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
von: Zheng, Ce, et al.
Veröffentlicht: (2026)
von: Zheng, Ce, et al.
Veröffentlicht: (2026)
SpecMemo: Speculative Decoding is in Your Pocket
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
von: Yildirim, Selin, et al.
Veröffentlicht: (2025)
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
CoCoI: Distributed Coded Inference System for Straggler Mitigation
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
PipeSpec: Breaking Stage Dependencies in Hierarchical LLM Decoding
von: McDanel, Bradley, et al.
Veröffentlicht: (2025)
von: McDanel, Bradley, et al.
Veröffentlicht: (2025)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
SPECTRE: Hybrid Ordinary-Parallel Speculative Serving for Resource-Efficient LLM Inference
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
von: Xie, Jincheng, et al.
Veröffentlicht: (2026)
SpecInF: Exploiting Idle GPU Resources in Distributed DL Training via Speculative Inference Filling
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
Staleness-Centric Optimizations for Parallel Diffusion MoE Inference
von: Luo, Jiajun, et al.
Veröffentlicht: (2024)
von: Luo, Jiajun, et al.
Veröffentlicht: (2024)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
von: Zhang, Han, et al.
Veröffentlicht: (2026)
von: Zhang, Han, et al.
Veröffentlicht: (2026)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
Accelerating OpenPangu Inference on NPU via Speculative Decoding
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
von: Dai, Yuntao, et al.
Veröffentlicht: (2026)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
von: Kumar, Satyam, et al.
Veröffentlicht: (2026)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
DynaKV: Enabling Accurate and Efficient Long-Sequence LLM Decoding on Smartphones
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
von: Wang, Tuowei, et al.
Veröffentlicht: (2025)
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
von: Abramovich, Talor, et al.
Veröffentlicht: (2026)
von: Abramovich, Talor, et al.
Veröffentlicht: (2026)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
von: Gond, Raja, et al.
Veröffentlicht: (2026)
von: Gond, Raja, et al.
Veröffentlicht: (2026)
Fast LLM Post-training via Decoupled and Fastest-of-N Speculation
von: Cheng, Rongxin, et al.
Veröffentlicht: (2025)
von: Cheng, Rongxin, et al.
Veröffentlicht: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
von: Xu, Chuhao, et al.
Veröffentlicht: (2025)
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
von: Xue, Zhenliang, et al.
Veröffentlicht: (2025)
von: Xue, Zhenliang, et al.
Veröffentlicht: (2025)
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
von: Wang, Shiju, et al.
Veröffentlicht: (2025)
SpecKV: Adaptive Speculative Decoding with Compression-Aware Gamma Selection
von: Shukla, Shikhar
Veröffentlicht: (2026)
von: Shukla, Shikhar
Veröffentlicht: (2026)
PIPO: Pipelined Offloading for Efficient Inference on Consumer Devices
von: Liu, Yangyijian, et al.
Veröffentlicht: (2025)
von: Liu, Yangyijian, et al.
Veröffentlicht: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
von: He, Yongchao, et al.
Veröffentlicht: (2025)
von: He, Yongchao, et al.
Veröffentlicht: (2025)
BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
von: Zheng, Wanyi, et al.
Veröffentlicht: (2025)
Accelerating End-Cloud Collaborative Inference via Near Bubble-free Pipeline Optimization
von: Gao, Luyao, et al.
Veröffentlicht: (2024)
von: Gao, Luyao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026) -
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission
von: Zheng, Ce, et al.
Veröffentlicht: (2026) -
SpecMemo: Speculative Decoding is in Your Pocket
von: Yildirim, Selin, et al.
Veröffentlicht: (2025) -
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
von: Xu, Jiaming, et al.
Veröffentlicht: (2025) -
CoCoI: Distributed Coded Inference System for Straggler Mitigation
von: Liu, Xing, et al.
Veröffentlicht: (2025)