DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Mingyu, Choi, Jae-Young, Moon, Kihyo, Jang, Minsung, Jeon, Eunjoo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
by: Khan, Kabir, et al.
Published: (2025)
by: Khan, Kabir, et al.
Published: (2025)
Knowledge Graphs-Driven Intelligence for Distributed Decision Systems
by: Napoli, Rosario, et al.
Published: (2026)
by: Napoli, Rosario, et al.
Published: (2026)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
by: Zehra, Sehar, et al.
Published: (2025)
by: Zehra, Sehar, et al.
Published: (2025)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
by: Li, Xiangchen, et al.
Published: (2026)
by: Li, Xiangchen, et al.
Published: (2026)
How Machine Learning-Data Driven Replication Strategies Enhance Fault Tolerance in Large-Scale Distributed Systems
by: Murimi, Almond Kiruthu
Published: (2025)
by: Murimi, Almond Kiruthu
Published: (2025)
StepCache: Step-Level Reuse with Lightweight Verification and Selective Patching for LLM Serving
by: Nouri, Azam
Published: (2026)
by: Nouri, Azam
Published: (2026)
ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor
by: Choi, Seungbeom, et al.
Published: (2025)
by: Choi, Seungbeom, et al.
Published: (2025)
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
by: Bin, Kyungmin, et al.
Published: (2025)
by: Bin, Kyungmin, et al.
Published: (2025)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
Experimentally Evaluating the Resource Efficiency of Big Data Autoscaling
by: Will, Jonathan, et al.
Published: (2025)
by: Will, Jonathan, et al.
Published: (2025)
DAGER: Exact Gradient Inversion for Large Language Models
by: Petrov, Ivo, et al.
Published: (2024)
by: Petrov, Ivo, et al.
Published: (2024)
Combining Serverless and High-Performance Computing Paradigms to support ML Data-Intensive Applications
by: Staylor, Mills, et al.
Published: (2025)
by: Staylor, Mills, et al.
Published: (2025)
Deep RC: A Scalable Data Engineering and Deep Learning Pipeline
by: Sarker, Arup Kumar, et al.
Published: (2025)
by: Sarker, Arup Kumar, et al.
Published: (2025)
Design and Implementation of an Analysis Pipeline for Heterogeneous Data
by: Sarker, Arup Kumar, et al.
Published: (2024)
by: Sarker, Arup Kumar, et al.
Published: (2024)
De-DSI: Decentralised Differentiable Search Index
by: Neague, Petru, et al.
Published: (2024)
by: Neague, Petru, et al.
Published: (2024)
Towards Message Brokers for Generative AI: Survey, Challenges, and Opportunities
by: Saleh, Alaa, et al.
Published: (2023)
by: Saleh, Alaa, et al.
Published: (2023)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
by: Kolluru, Saicharan
Published: (2025)
by: Kolluru, Saicharan
Published: (2025)
AAFLOW: Scalable Patterns for Agentic AI Workflows
by: Sarker, Arup Kumar, et al.
Published: (2026)
by: Sarker, Arup Kumar, et al.
Published: (2026)
CooperLLM: Cloud-Edge-End Cooperative Federated Fine-tuning for LLMs via ZOO-based Gradient Correction
by: Sun, He, et al.
Published: (2026)
by: Sun, He, et al.
Published: (2026)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
by: Kamath, Aditya K, et al.
Published: (2024)
by: Kamath, Aditya K, et al.
Published: (2024)
Flash-Fusion: Enabling Expressive, Low-Latency Queries on IoT Sensor Streams with LLMs
by: Patherya, Kausar, et al.
Published: (2025)
by: Patherya, Kausar, et al.
Published: (2025)
EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference
by: Sidik, Bronislav, et al.
Published: (2026)
by: Sidik, Bronislav, et al.
Published: (2026)
ACME: Adaptive Customization of Large Models via Distributed Systems
by: Dai, Ziming, et al.
Published: (2025)
by: Dai, Ziming, et al.
Published: (2025)
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
by: Li, Zonghang, et al.
Published: (2024)
by: Li, Zonghang, et al.
Published: (2024)
HFedATM: Hierarchical Federated Domain Generalization via Optimal Transport and Regularized Mean Aggregation
by: Nguyen, Thinh, et al.
Published: (2025)
by: Nguyen, Thinh, et al.
Published: (2025)
An Empirical Study of the Impact of Federated Learning on Machine Learning Model Accuracy
by: Yang, Haotian, et al.
Published: (2025)
by: Yang, Haotian, et al.
Published: (2025)
Deadline-Aware Joint Task Scheduling and Offloading in Mobile Edge Computing Systems
by: Nguyen, Ngoc Hung, et al.
Published: (2025)
by: Nguyen, Ngoc Hung, et al.
Published: (2025)
TAGC: Optimizing Gradient Communication in Distributed Transformer Training
by: Polyakov, Igor, et al.
Published: (2025)
by: Polyakov, Igor, et al.
Published: (2025)
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
by: Li, Zonghang, et al.
Published: (2025)
by: Li, Zonghang, et al.
Published: (2025)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
by: Penke, Carolin, et al.
Published: (2025)
by: Penke, Carolin, et al.
Published: (2025)
Kant: An Efficient Unified Scheduling System for Large-Scale AI Clusters
by: Zeng, Lingling, et al.
Published: (2025)
by: Zeng, Lingling, et al.
Published: (2025)
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
by: Zhang, Lingzhe, et al.
Published: (2025)
by: Zhang, Lingzhe, et al.
Published: (2025)
Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling
by: Luo, Zizhang, et al.
Published: (2026)
by: Luo, Zizhang, et al.
Published: (2026)
SLED: A Speculative LLM Decoding Framework for Efficient Edge Serving
by: Li, Xiangchen, et al.
Published: (2025)
by: Li, Xiangchen, et al.
Published: (2025)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
by: Chen, Mu-Chi, et al.
Published: (2025)
by: Chen, Mu-Chi, et al.
Published: (2025)
Augmenting the FedProx Algorithm by Minimizing Convergence
by: Sarkar, Anomitra, et al.
Published: (2024)
by: Sarkar, Anomitra, et al.
Published: (2024)
Shipwright: Proving liveness of distributed systems with Byzantine participants
by: Leung, Derek, et al.
Published: (2025)
by: Leung, Derek, et al.
Published: (2025)
Neural Router: Semantic Content Matching for Agentic AI
by: Lovén, Lauri, et al.
Published: (2026)
by: Lovén, Lauri, et al.
Published: (2026)
SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost
by: Li, Zhifei, et al.
Published: (2026)
by: Li, Zhifei, et al.
Published: (2026)
Similar Items
-
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
by: Khan, Kabir, et al.
Published: (2025) -
Knowledge Graphs-Driven Intelligence for Distributed Decision Systems
by: Napoli, Rosario, et al.
Published: (2026) -
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
by: Li, Xiangchen, et al.
Published: (2026) -
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
by: Zehra, Sehar, et al.
Published: (2025) -
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
by: Li, Xiangchen, et al.
Published: (2026)