DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Fengze, Li, Leshu, McDanel, Brad, Zhang, Sai Qian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
von: McDanel, Bradley
Veröffentlicht: (2024)
von: McDanel, Bradley
Veröffentlicht: (2024)
PipeSpec: Breaking Stage Dependencies in Hierarchical LLM Decoding
von: McDanel, Bradley, et al.
Veröffentlicht: (2025)
von: McDanel, Bradley, et al.
Veröffentlicht: (2025)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
von: Han, Yunhe, et al.
Veröffentlicht: (2026)
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
von: Tran, Phuong, et al.
Veröffentlicht: (2025)
Collaborative Speculative Inference for Efficient LLM Inference Serving
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
FlexSpec: Frozen Drafts Meet Evolving Targets in Edge-Cloud Collaborative LLM Speculative Decoding
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
FASER: Fine-Grained Phase Management for Speculative Decoding in Dynamic LLM Serving
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
von: Chen, Wenyan, et al.
Veröffentlicht: (2026)
SpecRouter: Adaptive Routing for Multi-Level Speculative Decoding in Large Language Models
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
von: Zhong, Yinmin, et al.
Veröffentlicht: (2024)
SwiftSpec: Ultra-Low Latency LLM Decoding by Scaling Asynchronous Speculative Decoding
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyi, et al.
Veröffentlicht: (2025)
Enabling Efficient Serverless Inference Serving for LLM (Large Language Model) in the Cloud
von: Ghosh, Himel
Veröffentlicht: (2024)
von: Ghosh, Himel
Veröffentlicht: (2024)
Distributed Speculative Execution for Resilient Cloud Applications
von: Li, Tianyu, et al.
Veröffentlicht: (2024)
von: Li, Tianyu, et al.
Veröffentlicht: (2024)
DLoRA: Distributed Parameter-Efficient Fine-Tuning Solution for Large Language Model
von: Gao, Chao, et al.
Veröffentlicht: (2024)
von: Gao, Chao, et al.
Veröffentlicht: (2024)
Stateful Large Language Model Serving with Pensieve
von: Yu, Lingfan, et al.
Veröffentlicht: (2023)
von: Yu, Lingfan, et al.
Veröffentlicht: (2023)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
von: Cao, Jiahe, et al.
Veröffentlicht: (2026)
von: Cao, Jiahe, et al.
Veröffentlicht: (2026)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
von: Xiang, Yuxing, et al.
Veröffentlicht: (2025)
von: Xiang, Yuxing, et al.
Veröffentlicht: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
von: Wei, Jinhui, et al.
Veröffentlicht: (2025)
EdgeServe: A Streaming System for Decentralized Model Serving
von: Shaowang, Ted, et al.
Veröffentlicht: (2023)
von: Shaowang, Ted, et al.
Veröffentlicht: (2023)
Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
von: Venkatesha, Yeshwanth, et al.
Veröffentlicht: (2025)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
von: Li, Zikun, et al.
Veröffentlicht: (2025)
von: Li, Zikun, et al.
Veröffentlicht: (2025)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2025)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
von: Zhan, Huiyou, et al.
Veröffentlicht: (2025)
von: Zhan, Huiyou, et al.
Veröffentlicht: (2025)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
von: Miao, Xupeng, et al.
Veröffentlicht: (2023)
MoLink: Distributed and Efficient Serving Framework for Large Models
von: Jin, Lewei, et al.
Veröffentlicht: (2025)
von: Jin, Lewei, et al.
Veröffentlicht: (2025)
Distributed Edge Analytics in Edge-Fog-Cloud Continuum
von: Srirama, Satish Narayana
Veröffentlicht: (2024)
von: Srirama, Satish Narayana
Veröffentlicht: (2024)
Intelligent Orchestration of Distributed Large Foundation Model Inference at the Edge
von: Koch, Fernando, et al.
Veröffentlicht: (2025)
von: Koch, Fernando, et al.
Veröffentlicht: (2025)
Towards Sustainable Large Language Model Serving
von: Nguyen, Sophia, et al.
Veröffentlicht: (2024)
von: Nguyen, Sophia, et al.
Veröffentlicht: (2024)
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
von: Wu, Bingyang, et al.
Veröffentlicht: (2024)
CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge Collaboration
von: Jin, Hongpeng, et al.
Veröffentlicht: (2024)
von: Jin, Hongpeng, et al.
Veröffentlicht: (2024)
Preble: Efficient Distributed Prompt Scheduling for LLM Serving
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2024)
von: Srivatsa, Vikranth, et al.
Veröffentlicht: (2024)
Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
von: Shi, Xiaoxiang, et al.
Veröffentlicht: (2025)
von: Shi, Xiaoxiang, et al.
Veröffentlicht: (2025)
Mélange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
von: Griggs, Tyler, et al.
Veröffentlicht: (2024)
von: Griggs, Tyler, et al.
Veröffentlicht: (2024)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
von: Chen, Xing, et al.
Veröffentlicht: (2025)
von: Chen, Xing, et al.
Veröffentlicht: (2025)
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2025)
AgileDART: An Agile and Scalable Edge Stream Processing Engine
von: Ching, Cheng-Wei, et al.
Veröffentlicht: (2024)
von: Ching, Cheng-Wei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM Acceleration
von: McDanel, Bradley
Veröffentlicht: (2024) -
PipeSpec: Breaking Stage Dependencies in Hierarchical LLM Decoding
von: McDanel, Bradley, et al.
Veröffentlicht: (2025) -
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026) -
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
von: Han, Yunhe, et al.
Veröffentlicht: (2026) -
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
von: Li, Rui, et al.
Veröffentlicht: (2025)