DiP-SD: Distributed Pipelined Speculative Decoding for Efficient LLM Inference at the Edge
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yaodan, Zhou, Sheng, Niu, Zhisheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WISV: Wireless-Informed Semantic Verification for Distributed Speculative Decoding in Device-Edge LLM Inference
by: Liu, Zixuan, et al.
Published: (2026)
by: Liu, Zixuan, et al.
Published: (2026)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
by: Xu, Yaodan, et al.
Published: (2025)
by: Xu, Yaodan, et al.
Published: (2025)
Conformal Sparsification for Bandwidth-Efficient Edge-Cloud Speculative Decoding
by: Bhattacharjee, Payel, et al.
Published: (2025)
by: Bhattacharjee, Payel, et al.
Published: (2025)
Robust DNN Partitioning and Resource Allocation Under Uncertain Inference Time
by: Nan, Zhaojun, et al.
Published: (2025)
by: Nan, Zhaojun, et al.
Published: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
by: Han, Yunhe, et al.
Published: (2026)
by: Han, Yunhe, et al.
Published: (2026)
SMDP-Based Dynamic Batching for Improving Responsiveness and Energy Efficiency of Batch Services
by: Xu, Yaodan, et al.
Published: (2025)
by: Xu, Yaodan, et al.
Published: (2025)
TGPP: Trajectory-Guided Plug-and-Play Priors for Sparse Radio Map Reconstruction
by: Zhang, Jiawen, et al.
Published: (2026)
by: Zhang, Jiawen, et al.
Published: (2026)
Reconfigurable Intelligent Surface for Green Edge Inference
by: Hua, Sheng, et al.
Published: (2019)
by: Hua, Sheng, et al.
Published: (2019)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
by: Liu, Xing, et al.
Published: (2025)
by: Liu, Xing, et al.
Published: (2025)
GELATO: Generative Entropy- and Lyapunov-based Adaptive Token Offloading for Device-Edge Speculative LLM Inference
by: Tang, Zengzipeng, et al.
Published: (2026)
by: Tang, Zengzipeng, et al.
Published: (2026)
Energy-Efficient Edge Inference in Integrated Sensing, Communication, and Computation Networks
by: Yao, Jiacheng, et al.
Published: (2025)
by: Yao, Jiacheng, et al.
Published: (2025)
Task-Oriented Wireless Communications for Collaborative Perception in Intelligent Unmanned Systems
by: Zhou, Sheng, et al.
Published: (2024)
by: Zhou, Sheng, et al.
Published: (2024)
Speeding up Speculative Decoding via Sequential Approximate Verification
by: Zhong, Meiyu, et al.
Published: (2025)
by: Zhong, Meiyu, et al.
Published: (2025)
DiP: Taming Diffusion Models in Pixel Space
by: Chen, Zhennan, et al.
Published: (2025)
by: Chen, Zhennan, et al.
Published: (2025)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
by: Zhang, Yida, et al.
Published: (2026)
by: Zhang, Yida, et al.
Published: (2026)
Greedy Multi-Path Block Verification for Faster Decoding in Speculative Sampling
by: Thomas, Rahul, et al.
Published: (2026)
by: Thomas, Rahul, et al.
Published: (2026)
C-MASS: Combinatorial Mobility-Aware Sensor Scheduling for Collaborative Perception with Second-Order Topology Approximation
by: Jia, Yukuan, et al.
Published: (2024)
by: Jia, Yukuan, et al.
Published: (2024)
Block Verification Accelerates Speculative Decoding
by: Sun, Ziteng, et al.
Published: (2024)
by: Sun, Ziteng, et al.
Published: (2024)
Efficiency Unleashed: Inference Acceleration for LLM-based Recommender Systems with Speculative Decoding
by: Xi, Yunjia, et al.
Published: (2024)
by: Xi, Yunjia, et al.
Published: (2024)
Channel Capacity-Aware Distributed Encoding for Multi-View Sensing and Edge Inference
by: Yang, Mingjie, et al.
Published: (2024)
by: Yang, Mingjie, et al.
Published: (2024)
METTLE: Efficient Streaming Erasure Code with Peeling Decodability
by: Yu, Qianru, et al.
Published: (2026)
by: Yu, Qianru, et al.
Published: (2026)
Sparse Optimization for Green Edge AI Inference
by: Yang, Xiangyu, et al.
Published: (2020)
by: Yang, Xiangyu, et al.
Published: (2020)
Error Exponents for Randomised List Decoding
by: Miyamoto, Henrique K., et al.
Published: (2026)
by: Miyamoto, Henrique K., et al.
Published: (2026)
Speculative Decoding Scaling Laws (SDSL): Throughput Optimization Made Simple
by: Bozorgkhoo, Amirhossein, et al.
Published: (2026)
by: Bozorgkhoo, Amirhossein, et al.
Published: (2026)
Future Validity is the Missing Statistic: From Impossibility to $Φ$-Estimation for Grammar-Faithful Speculative Decoding
by: Nie, Wenhua, et al.
Published: (2026)
by: Nie, Wenhua, et al.
Published: (2026)
Dynamic Scheduling for Vehicle-to-Vehicle Communications Enhanced Federated Learning
by: Yan, Jintao, et al.
Published: (2024)
by: Yan, Jintao, et al.
Published: (2024)
Unrolled and Pipelined Decoders based on Look-Up Tables for Polar Codes
by: Giard, Pascal, et al.
Published: (2023)
by: Giard, Pascal, et al.
Published: (2023)
CR^2: Cost-Aware Risk-Controlled Routing for Wireless Device-Edge LLM Inference
by: Xue, Nan, et al.
Published: (2026)
by: Xue, Nan, et al.
Published: (2026)
AdaSD: Adaptive Speculative Decoding for Efficient Language Model Inference
by: Lu, Kuan-Wei, et al.
Published: (2025)
by: Lu, Kuan-Wei, et al.
Published: (2025)
Edge-Based Anisotropic Decoding for Generalized Bicycle Codes
by: Chytas, Dimitris, et al.
Published: (2026)
by: Chytas, Dimitris, et al.
Published: (2026)
Blind Recognition of Polar Codes Using Successive Cancellation List Decoding
by: Tu, Changwei, et al.
Published: (2026)
by: Tu, Changwei, et al.
Published: (2026)
Distributed Indirect Source Coding with Decoder Side Information
by: Tang, Jiancheng, et al.
Published: (2024)
by: Tang, Jiancheng, et al.
Published: (2024)
Universal Decoding over Finite-State Additive Channels via Noise Guessing
by: Miyamoto, Henrique K., et al.
Published: (2025)
by: Miyamoto, Henrique K., et al.
Published: (2025)
SpecTr: Fast Speculative Decoding via Optimal Transport
by: Sun, Ziteng, et al.
Published: (2023)
by: Sun, Ziteng, et al.
Published: (2023)
DiP: A Scalable, Energy-Efficient Systolic Array for Matrix Multiplication Acceleration
by: Abdelmaksoud, Ahmed J., et al.
Published: (2024)
by: Abdelmaksoud, Ahmed J., et al.
Published: (2024)
Dynamic Quality-Latency Aware Routing for LLM Inference in Wireless Edge-Device Networks
by: Bao, Rui, et al.
Published: (2025)
by: Bao, Rui, et al.
Published: (2025)
Analysis of Efficient Scheduling in Layered Decoding of GLDPC Codes
by: Peng, Qingqing, et al.
Published: (2026)
by: Peng, Qingqing, et al.
Published: (2026)
Efficient LLR-Domain Decoding of ABS+ Polar Codes
by: Chernikov, Mikhail, et al.
Published: (2026)
by: Chernikov, Mikhail, et al.
Published: (2026)
Graphical Models and Efficient Inference Methods for Multivariate Phase Probability Distributions
by: Perley, Andrew S., et al.
Published: (2025)
by: Perley, Andrew S., et al.
Published: (2025)
Edge Perception: Intelligent Wireless Sensing at Network Edge
by: Cui, Yuanhao, et al.
Published: (2024)
by: Cui, Yuanhao, et al.
Published: (2024)
Similar Items
-
WISV: Wireless-Informed Semantic Verification for Distributed Speculative Decoding in Device-Edge LLM Inference
by: Liu, Zixuan, et al.
Published: (2026) -
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
by: Xu, Yaodan, et al.
Published: (2025) -
Conformal Sparsification for Bandwidth-Efficient Edge-Cloud Speculative Decoding
by: Bhattacharjee, Payel, et al.
Published: (2025) -
Robust DNN Partitioning and Resource Allocation Under Uncertain Inference Time
by: Nan, Zhaojun, et al.
Published: (2025) -
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
by: Han, Yunhe, et al.
Published: (2026)