DiP-SD: Distributed Pipelined Speculative Decoding for Efficient LLM Inference at the Edge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Yaodan, Zhou, Sheng, Niu, Zhisheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WISV: Wireless-Informed Semantic Verification for Distributed Speculative Decoding in Device-Edge LLM Inference
von: Liu, Zixuan, et al.
Veröffentlicht: (2026)
von: Liu, Zixuan, et al.
Veröffentlicht: (2026)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
von: Xu, Yaodan, et al.
Veröffentlicht: (2025)
von: Xu, Yaodan, et al.
Veröffentlicht: (2025)
Conformal Sparsification for Bandwidth-Efficient Edge-Cloud Speculative Decoding
von: Bhattacharjee, Payel, et al.
Veröffentlicht: (2025)
von: Bhattacharjee, Payel, et al.
Veröffentlicht: (2025)
Robust DNN Partitioning and Resource Allocation Under Uncertain Inference Time
von: Nan, Zhaojun, et al.
Veröffentlicht: (2025)
von: Nan, Zhaojun, et al.
Veröffentlicht: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
SMDP-Based Dynamic Batching for Improving Responsiveness and Energy Efficiency of Batch Services
von: Xu, Yaodan, et al.
Veröffentlicht: (2025)
von: Xu, Yaodan, et al.
Veröffentlicht: (2025)
TGPP: Trajectory-Guided Plug-and-Play Priors for Sparse Radio Map Reconstruction
von: Zhang, Jiawen, et al.
Veröffentlicht: (2026)
von: Zhang, Jiawen, et al.
Veröffentlicht: (2026)
Reconfigurable Intelligent Surface for Green Edge Inference
von: Hua, Sheng, et al.
Veröffentlicht: (2019)
von: Hua, Sheng, et al.
Veröffentlicht: (2019)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
GELATO: Generative Entropy- and Lyapunov-based Adaptive Token Offloading for Device-Edge Speculative LLM Inference
von: Tang, Zengzipeng, et al.
Veröffentlicht: (2026)
von: Tang, Zengzipeng, et al.
Veröffentlicht: (2026)
Energy-Efficient Edge Inference in Integrated Sensing, Communication, and Computation Networks
von: Yao, Jiacheng, et al.
Veröffentlicht: (2025)
von: Yao, Jiacheng, et al.
Veröffentlicht: (2025)
Task-Oriented Wireless Communications for Collaborative Perception in Intelligent Unmanned Systems
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
Speeding up Speculative Decoding via Sequential Approximate Verification
von: Zhong, Meiyu, et al.
Veröffentlicht: (2025)
von: Zhong, Meiyu, et al.
Veröffentlicht: (2025)
DiP: Taming Diffusion Models in Pixel Space
von: Chen, Zhennan, et al.
Veröffentlicht: (2025)
von: Chen, Zhennan, et al.
Veröffentlicht: (2025)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
Greedy Multi-Path Block Verification for Faster Decoding in Speculative Sampling
von: Thomas, Rahul, et al.
Veröffentlicht: (2026)
von: Thomas, Rahul, et al.
Veröffentlicht: (2026)
C-MASS: Combinatorial Mobility-Aware Sensor Scheduling for Collaborative Perception with Second-Order Topology Approximation
von: Jia, Yukuan, et al.
Veröffentlicht: (2024)
von: Jia, Yukuan, et al.
Veröffentlicht: (2024)
Block Verification Accelerates Speculative Decoding
von: Sun, Ziteng, et al.
Veröffentlicht: (2024)
von: Sun, Ziteng, et al.
Veröffentlicht: (2024)
Efficiency Unleashed: Inference Acceleration for LLM-based Recommender Systems with Speculative Decoding
von: Xi, Yunjia, et al.
Veröffentlicht: (2024)
von: Xi, Yunjia, et al.
Veröffentlicht: (2024)
Channel Capacity-Aware Distributed Encoding for Multi-View Sensing and Edge Inference
von: Yang, Mingjie, et al.
Veröffentlicht: (2024)
von: Yang, Mingjie, et al.
Veröffentlicht: (2024)
METTLE: Efficient Streaming Erasure Code with Peeling Decodability
von: Yu, Qianru, et al.
Veröffentlicht: (2026)
von: Yu, Qianru, et al.
Veröffentlicht: (2026)
Sparse Optimization for Green Edge AI Inference
von: Yang, Xiangyu, et al.
Veröffentlicht: (2020)
von: Yang, Xiangyu, et al.
Veröffentlicht: (2020)
Error Exponents for Randomised List Decoding
von: Miyamoto, Henrique K., et al.
Veröffentlicht: (2026)
von: Miyamoto, Henrique K., et al.
Veröffentlicht: (2026)
Speculative Decoding Scaling Laws (SDSL): Throughput Optimization Made Simple
von: Bozorgkhoo, Amirhossein, et al.
Veröffentlicht: (2026)
von: Bozorgkhoo, Amirhossein, et al.
Veröffentlicht: (2026)
Future Validity is the Missing Statistic: From Impossibility to $Φ$-Estimation for Grammar-Faithful Speculative Decoding
von: Nie, Wenhua, et al.
Veröffentlicht: (2026)
von: Nie, Wenhua, et al.
Veröffentlicht: (2026)
Dynamic Scheduling for Vehicle-to-Vehicle Communications Enhanced Federated Learning
von: Yan, Jintao, et al.
Veröffentlicht: (2024)
von: Yan, Jintao, et al.
Veröffentlicht: (2024)
Unrolled and Pipelined Decoders based on Look-Up Tables for Polar Codes
von: Giard, Pascal, et al.
Veröffentlicht: (2023)
von: Giard, Pascal, et al.
Veröffentlicht: (2023)
CR^2: Cost-Aware Risk-Controlled Routing for Wireless Device-Edge LLM Inference
von: Xue, Nan, et al.
Veröffentlicht: (2026)
von: Xue, Nan, et al.
Veröffentlicht: (2026)
AdaSD: Adaptive Speculative Decoding for Efficient Language Model Inference
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
Edge-Based Anisotropic Decoding for Generalized Bicycle Codes
von: Chytas, Dimitris, et al.
Veröffentlicht: (2026)
von: Chytas, Dimitris, et al.
Veröffentlicht: (2026)
Blind Recognition of Polar Codes Using Successive Cancellation List Decoding
von: Tu, Changwei, et al.
Veröffentlicht: (2026)
von: Tu, Changwei, et al.
Veröffentlicht: (2026)
Distributed Indirect Source Coding with Decoder Side Information
von: Tang, Jiancheng, et al.
Veröffentlicht: (2024)
von: Tang, Jiancheng, et al.
Veröffentlicht: (2024)
Universal Decoding over Finite-State Additive Channels via Noise Guessing
von: Miyamoto, Henrique K., et al.
Veröffentlicht: (2025)
von: Miyamoto, Henrique K., et al.
Veröffentlicht: (2025)
SpecTr: Fast Speculative Decoding via Optimal Transport
von: Sun, Ziteng, et al.
Veröffentlicht: (2023)
von: Sun, Ziteng, et al.
Veröffentlicht: (2023)
DiP: A Scalable, Energy-Efficient Systolic Array for Matrix Multiplication Acceleration
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2024)
von: Abdelmaksoud, Ahmed J., et al.
Veröffentlicht: (2024)
Dynamic Quality-Latency Aware Routing for LLM Inference in Wireless Edge-Device Networks
von: Bao, Rui, et al.
Veröffentlicht: (2025)
von: Bao, Rui, et al.
Veröffentlicht: (2025)
Analysis of Efficient Scheduling in Layered Decoding of GLDPC Codes
von: Peng, Qingqing, et al.
Veröffentlicht: (2026)
von: Peng, Qingqing, et al.
Veröffentlicht: (2026)
Efficient LLR-Domain Decoding of ABS+ Polar Codes
von: Chernikov, Mikhail, et al.
Veröffentlicht: (2026)
von: Chernikov, Mikhail, et al.
Veröffentlicht: (2026)
Graphical Models and Efficient Inference Methods for Multivariate Phase Probability Distributions
von: Perley, Andrew S., et al.
Veröffentlicht: (2025)
von: Perley, Andrew S., et al.
Veröffentlicht: (2025)
Edge Perception: Intelligent Wireless Sensing at Network Edge
von: Cui, Yuanhao, et al.
Veröffentlicht: (2024)
von: Cui, Yuanhao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
WISV: Wireless-Informed Semantic Verification for Distributed Speculative Decoding in Device-Edge LLM Inference
von: Liu, Zixuan, et al.
Veröffentlicht: (2026) -
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
von: Xu, Yaodan, et al.
Veröffentlicht: (2025) -
Conformal Sparsification for Bandwidth-Efficient Edge-Cloud Speculative Decoding
von: Bhattacharjee, Payel, et al.
Veröffentlicht: (2025) -
Robust DNN Partitioning and Resource Allocation Under Uncertain Inference Time
von: Nan, Zhaojun, et al.
Veröffentlicht: (2025) -
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026)