PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Yunhe, Gao, Yunqi, Hu, Bing, Mashhadi, Mahdi Boloursaz, Duan, Yitong, Xiao, Pei, Zhang, Yanfeng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918520723341312
author Han, Yunhe
Gao, Yunqi
Hu, Bing
Mashhadi, Mahdi Boloursaz
Duan, Yitong
Xiao, Pei
Zhang, Yanfeng
author_facet Han, Yunhe
Gao, Yunqi
Hu, Bing
Mashhadi, Mahdi Boloursaz
Duan, Yitong
Xiao, Pei
Zhang, Yanfeng
contents Speculative decoding can significantly accelerate LLM inference, especially given that its cloud-edge collaborative deployment offers cloud workload offloading, offline robustness, and privacy enhancement. However, existing collaborative inference frameworks with speculative decoding are constrained by (i) sequential token generation and communication with low resource utilization, and (ii) inflexible cloud non-autoregressive verification (NAV) triggering that induces premature verification or costly rollbacks. In this paper, we propose PipeSD, an efficient cloud-edge collaborative pipeline inference framework with speculative decoding. PipeSD overlaps token generation and communication by a token-batch pipeline scheduling mechanism optimized by dynamic programming, and improves verification flexibility through a dual-threshold NAV triggering mechanism with a lightweight Bayesian optimization autotuner. We implement PipeSD using llama-cpp-python, PyTorch, and FastAPI, and evaluate it on a real-world cloud-edge testbed with two draft-target model pairs across four scenarios. Results show that PipeSD consistently outperforms state-of-the-art baselines, achieving 1.16x-2.16x speedup and reducing energy consumption by 14.3%-25.3%.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13319
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
Han, Yunhe
Gao, Yunqi
Hu, Bing
Mashhadi, Mahdi Boloursaz
Duan, Yitong
Xiao, Pei
Zhang, Yanfeng
Distributed, Parallel, and Cluster Computing
Speculative decoding can significantly accelerate LLM inference, especially given that its cloud-edge collaborative deployment offers cloud workload offloading, offline robustness, and privacy enhancement. However, existing collaborative inference frameworks with speculative decoding are constrained by (i) sequential token generation and communication with low resource utilization, and (ii) inflexible cloud non-autoregressive verification (NAV) triggering that induces premature verification or costly rollbacks. In this paper, we propose PipeSD, an efficient cloud-edge collaborative pipeline inference framework with speculative decoding. PipeSD overlaps token generation and communication by a token-batch pipeline scheduling mechanism optimized by dynamic programming, and improves verification flexibility through a dual-threshold NAV triggering mechanism with a lightweight Bayesian optimization autotuner. We implement PipeSD using llama-cpp-python, PyTorch, and FastAPI, and evaluate it on a real-world cloud-edge testbed with two draft-target model pairs across four scenarios. Results show that PipeSD consistently outperforms state-of-the-art baselines, achieving 1.16x-2.16x speedup and reducing energy consumption by 14.3%-25.3%.
title PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2605.13319