SLED: A Speculative LLM Decoding Framework for Efficient Edge Serving
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Xiangchen, Spatharakis, Dimitrios, Ghafouri, Saeid, Fan, Jiakun, Vandierendonck, Hans, John, Deepu, Ji, Bo, Nikolopoulos, Dimitrios |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
Tetris: An SLA-aware Application Placement Strategy in the Edge-Cloud Continuum
von: Almeida, Lucas, et al.
Veröffentlicht: (2025)
von: Almeida, Lucas, et al.
Veröffentlicht: (2025)
A Comprehensive Survey on Orbital Edge Computing: Systems, Applications, and Algorithms
von: Wu, Changhao, et al.
Veröffentlicht: (2023)
von: Wu, Changhao, et al.
Veröffentlicht: (2023)
Rendezvous and Merging for Two Metamorphic Robotic Systems without Global Compass
von: Yamada, Ryonosuke, et al.
Veröffentlicht: (2024)
von: Yamada, Ryonosuke, et al.
Veröffentlicht: (2024)
Implementation and Evaluation of Fast Raft for Hierarchical Consensus
von: Melnychuk, Anton, et al.
Veröffentlicht: (2025)
von: Melnychuk, Anton, et al.
Veröffentlicht: (2025)
Light Cone Consistency: Toward a Unified Theory of Consistency in Message-Passing Systems
von: Landers, Rob, et al.
Veröffentlicht: (2026)
von: Landers, Rob, et al.
Veröffentlicht: (2026)
Accelerating Frontier MoE Training with 3D Integrated Optics
von: Bernadskiy, Mikhail, et al.
Veröffentlicht: (2025)
von: Bernadskiy, Mikhail, et al.
Veröffentlicht: (2025)
CRDT-Based Game State Synchronization in Peer-to-Peer VR
von: Dantas, Abel, et al.
Veröffentlicht: (2025)
von: Dantas, Abel, et al.
Veröffentlicht: (2025)
LLM Agents for Interactive Workflow Provenance: Reference Architecture and Evaluation Methodology
von: Souza, Renan, et al.
Veröffentlicht: (2025)
von: Souza, Renan, et al.
Veröffentlicht: (2025)
Splitwise: Collaborative Edge-Cloud Inference for LLMs via Lyapunov-Assisted DRL
von: Younesi, Abolfazl, et al.
Veröffentlicht: (2025)
von: Younesi, Abolfazl, et al.
Veröffentlicht: (2025)
Next-Generation Event-Driven Architectures: Performance, Scalability, and Intelligent Orchestration Across Messaging Frameworks
von: Arafat, Jahidul, et al.
Veröffentlicht: (2025)
von: Arafat, Jahidul, et al.
Veröffentlicht: (2025)
Optimizing Multi-DNN Inference on Mobile Devices through Heterogeneous Processor Co-Execution
von: Gao, Yunquan, et al.
Veröffentlicht: (2025)
von: Gao, Yunquan, et al.
Veröffentlicht: (2025)
Reexamining Paradigms of End-to-End Data Movement
von: Fang, Chin, et al.
Veröffentlicht: (2025)
von: Fang, Chin, et al.
Veröffentlicht: (2025)
NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
von: Goldman, Amos, et al.
Veröffentlicht: (2026)
von: Goldman, Amos, et al.
Veröffentlicht: (2026)
Representation Similarity: A Better Guidance of DNN Layer Sharing for Edge Computing without Training
von: Cao, Bryan Bo, et al.
Veröffentlicht: (2024)
von: Cao, Bryan Bo, et al.
Veröffentlicht: (2024)
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling
von: Li, Yinrong, et al.
Veröffentlicht: (2026)
von: Li, Yinrong, et al.
Veröffentlicht: (2026)
PROV-AGENT: Unified Provenance for Tracking AI Agent Interactions in Agentic Workflows
von: Souza, Renan, et al.
Veröffentlicht: (2025)
von: Souza, Renan, et al.
Veröffentlicht: (2025)
Model-driven development of data intensive applications over cloud resources
von: Tolosana-Calasanz, Rafael, et al.
Veröffentlicht: (2024)
von: Tolosana-Calasanz, Rafael, et al.
Veröffentlicht: (2024)
Efficiently Scheduling Parallel DAG Tasks on Identical Multiprocessors
von: Lendve, Shardul, et al.
Veröffentlicht: (2024)
von: Lendve, Shardul, et al.
Veröffentlicht: (2024)
GPU-Initiated Networking for NCCL
von: Hamidouche, Khaled, et al.
Veröffentlicht: (2025)
von: Hamidouche, Khaled, et al.
Veröffentlicht: (2025)
Coordinated Reinforcement Learning Prefetching Architecture for Multicore Systems
von: Siddiqui, Mohammed Humaid, et al.
Veröffentlicht: (2025)
von: Siddiqui, Mohammed Humaid, et al.
Veröffentlicht: (2025)
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
von: Li, Zonghang, et al.
Veröffentlicht: (2024)
von: Li, Zonghang, et al.
Veröffentlicht: (2024)
Replication in Graph Partitioning and Scheduling Problems
von: Papp, Pál András, et al.
Veröffentlicht: (2026)
von: Papp, Pál András, et al.
Veröffentlicht: (2026)
FlashSpread: IO-Aware GPU Simulation of Non-Markovian Epidemic Dynamics via Kernel Fusion
von: Shakeri, Heman, et al.
Veröffentlicht: (2026)
von: Shakeri, Heman, et al.
Veröffentlicht: (2026)
The $qs$ Inequality: Quantifying the Double Penalty of Mixture-of-Experts at Inference
von: Adhinarayanan, Vignesh, et al.
Veröffentlicht: (2026)
von: Adhinarayanan, Vignesh, et al.
Veröffentlicht: (2026)
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal Correlation
von: Cheng, Long, et al.
Veröffentlicht: (2026)
von: Cheng, Long, et al.
Veröffentlicht: (2026)
Lincoln AI Computing Survey (LAICS) and Trends
von: Reuther, Albert, et al.
Veröffentlicht: (2025)
von: Reuther, Albert, et al.
Veröffentlicht: (2025)
Directives for Function Offloading in 5G Networks Based on a Performance Characteristics Analysis
von: Dettinger, Falk, et al.
Veröffentlicht: (2025)
von: Dettinger, Falk, et al.
Veröffentlicht: (2025)
Efficient Multi-Processor Scheduling in Increasingly Realistic Models
von: Papp, Pál András, et al.
Veröffentlicht: (2024)
von: Papp, Pál András, et al.
Veröffentlicht: (2024)
Cognitive Infrastructure: A Unified DCIM Framework for AI Data Centers
von: Sunkara, Krishna Chaitanya
Veröffentlicht: (2026)
von: Sunkara, Krishna Chaitanya
Veröffentlicht: (2026)
SCION: Size-aware Policy Orchestration for Nonstationary Object Caches (Long Paper Version)
von: Wang, Qizhi
Veröffentlicht: (2026)
von: Wang, Qizhi
Veröffentlicht: (2026)
In search of the lost tree: Hardness and relaxation of spanning trees in temporal graphs
von: Casteigts, Arnaud, et al.
Veröffentlicht: (2023)
von: Casteigts, Arnaud, et al.
Veröffentlicht: (2023)
Simple, strict, proper, happy: A study of reachability in temporal graphs
von: Casteigts, Arnaud, et al.
Veröffentlicht: (2022)
von: Casteigts, Arnaud, et al.
Veröffentlicht: (2022)
Federated Learning in Adversarial Environments: Testbed Design and Poisoning Resilience in Cybersecurity
von: Huang, Hao Jian, et al.
Veröffentlicht: (2024)
von: Huang, Hao Jian, et al.
Veröffentlicht: (2024)
Moonshot: Optimizing Chain-Based Rotating Leader BFT via Optimistic Proposals
von: Doidge, Isaac, et al.
Veröffentlicht: (2024)
von: Doidge, Isaac, et al.
Veröffentlicht: (2024)
LAMMPS-KOKKOS: Performance Portable Molecular Dynamics Across Exascale Architectures
von: Johansson, Anders, et al.
Veröffentlicht: (2025)
von: Johansson, Anders, et al.
Veröffentlicht: (2025)
FPGA-Accelerated Lock Management and Transaction Processing: Architecture, Optimization, and Design Space Exploration
von: Zhu, Shien, et al.
Veröffentlicht: (2026)
von: Zhu, Shien, et al.
Veröffentlicht: (2026)
A Bio-Inspired Leader-based Energy Management System for Drone Fleets
von: Napoli, Rosario, et al.
Veröffentlicht: (2025)
von: Napoli, Rosario, et al.
Veröffentlicht: (2025)
Bridging Generalization Gap of Heterogeneous Federated Clients Using Generative Models
von: Niu, Ziru, et al.
Veröffentlicht: (2025)
von: Niu, Ziru, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
von: Li, Xiangchen, et al.
Veröffentlicht: (2026) -
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
von: Li, Xiangchen, et al.
Veröffentlicht: (2026) -
Tetris: An SLA-aware Application Placement Strategy in the Edge-Cloud Continuum
von: Almeida, Lucas, et al.
Veröffentlicht: (2025) -
A Comprehensive Survey on Orbital Edge Computing: Systems, Applications, and Algorithms
von: Wu, Changhao, et al.
Veröffentlicht: (2023) -
Rendezvous and Merging for Two Metamorphic Robotic Systems without Global Compass
von: Yamada, Ryonosuke, et al.
Veröffentlicht: (2024)