WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Xiangchen, Fan, Jiakun, Wang, Qingyuan, Spatharakis, Dimitrios, Ghafouri, Saeid, Vandierendonck, Hans, John, Deepu, Ji, Bo, Butt, Ali R., Nikolopoulos, Dimitrios S. |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
par: Li, Xiangchen, et autres
Publié: (2026)
par: Li, Xiangchen, et autres
Publié: (2026)
SLED: A Speculative LLM Decoding Framework for Efficient Edge Serving
par: Li, Xiangchen, et autres
Publié: (2025)
par: Li, Xiangchen, et autres
Publié: (2025)
ICAN-Deploy: Identity-Stable Canary Deployment for Safety-Critical Embodied Agents
par: Qin, Xue, et autres
Publié: (2026)
par: Qin, Xue, et autres
Publié: (2026)
Secure coding for web applications: Frameworks, challenges, and the role of LLMs
par: Kiashemshaki, Kiana, et autres
Publié: (2025)
par: Kiashemshaki, Kiana, et autres
Publié: (2025)
Secure and Scalable Blockchain Voting: A Comparative Framework and the Role of Large Language Models
par: Kiashemshaki, Kiana, et autres
Publié: (2025)
par: Kiashemshaki, Kiana, et autres
Publié: (2025)
A History Equivalence Algorithm for Dynamic Process Migration
par: Bakshi, Gargi, et autres
Publié: (2024)
par: Bakshi, Gargi, et autres
Publié: (2024)
Governed Capability Evolution: Lifecycle-Time Compatibility Checking and Rollback for AI-Component-Based Systems, with Embodied Agents as Case Study
par: Qin, Xue, et autres
Publié: (2026)
par: Qin, Xue, et autres
Publié: (2026)
LLMs as Idiomatic Decompilers: Recovering High-Level Code from x86-64 Assembly for Dart
par: Abualazm, Raafat, et autres
Publié: (2026)
par: Abualazm, Raafat, et autres
Publié: (2026)
C8s: A Confidential Kubernetes Architecture
par: Asad, Amean, et autres
Publié: (2026)
par: Asad, Amean, et autres
Publié: (2026)
Speculative Automated Refactoring of Imperative Deep Learning Programs to Graph Execution
par: Khatchadourian, Raffi, et autres
Publié: (2025)
par: Khatchadourian, Raffi, et autres
Publié: (2025)
Predicting Known Vulnerabilities from Attack Descriptions Using Sentence Transformers
par: Othman, Refat
Publié: (2026)
par: Othman, Refat
Publié: (2026)
Analyzing the Adoption of Database Management Systems Throughout the History of Open Source Projects
par: Paiva, Camila A., et autres
Publié: (2026)
par: Paiva, Camila A., et autres
Publié: (2026)
Site Reliability Engineering (SRE) and Observations on SRE Process to Make Tasks Easier
par: Puli, Balaram
Publié: (2025)
par: Puli, Balaram
Publié: (2025)
Semantic Analysis of Macro Usage for Portability
par: Pappas, Brent, et autres
Publié: (2024)
par: Pappas, Brent, et autres
Publié: (2024)
LeanBin: Harnessing Lifting and Recompilation to Debloat Binaries
par: Wodiany, Igor, et autres
Publié: (2024)
par: Wodiany, Igor, et autres
Publié: (2024)
Sink or SWIM: Tackling Real-Time ASR at Scale
par: Bruzzone, Federico, et autres
Publié: (2026)
par: Bruzzone, Federico, et autres
Publié: (2026)
Finding a Crab in the C: Assured Translation via Comparative Symbolic Execution
par: Helbling, Caleb, et autres
Publié: (2026)
par: Helbling, Caleb, et autres
Publié: (2026)
cozy: Comparative Symbolic Execution for Binary Programs
par: Helbling, Caleb, et autres
Publié: (2025)
par: Helbling, Caleb, et autres
Publié: (2025)
Combining Serverless and High-Performance Computing Paradigms to support ML Data-Intensive Applications
par: Staylor, Mills, et autres
Publié: (2025)
par: Staylor, Mills, et autres
Publié: (2025)
Deep RC: A Scalable Data Engineering and Deep Learning Pipeline
par: Sarker, Arup Kumar, et autres
Publié: (2025)
par: Sarker, Arup Kumar, et autres
Publié: (2025)
Design and Implementation of an Analysis Pipeline for Heterogeneous Data
par: Sarker, Arup Kumar, et autres
Publié: (2024)
par: Sarker, Arup Kumar, et autres
Publié: (2024)
Towards Observation Lakehouses: Living, Interactive Archives of Software Behavior
par: Kessel, Marcus
Publié: (2025)
par: Kessel, Marcus
Publié: (2025)
Is (Selective) Round-To-Nearest Quantization All You Need?
par: Kogan, Alex
Publié: (2025)
par: Kogan, Alex
Publié: (2025)
DPDPU: Data Processing with DPUs
par: Hu, Jiasheng, et autres
Publié: (2024)
par: Hu, Jiasheng, et autres
Publié: (2024)
Context-Instrumental Data Distillation for Kubernetes Manifest Generation: Method and Experimental Evaluation
par: Kozachok, Andrey, et autres
Publié: (2026)
par: Kozachok, Andrey, et autres
Publié: (2026)
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking
par: Guo, Dongxin, et autres
Publié: (2026)
par: Guo, Dongxin, et autres
Publié: (2026)
Towards a Probabilistic Framework for Analyzing and Improving LLM-Enabled Software
par: Baldonado, Juan Manuel, et autres
Publié: (2025)
par: Baldonado, Juan Manuel, et autres
Publié: (2025)
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
par: Lucas, Tom, et autres
Publié: (2026)
par: Lucas, Tom, et autres
Publié: (2026)
PoCL-R: An Open Standard Based Offloading Layer for Heterogeneous Multi-Access Edge Computing with Server Side Scalability
par: Solanti, Jan, et autres
Publié: (2023)
par: Solanti, Jan, et autres
Publié: (2023)
Spreadsheet Engineering: A Research Framework
par: Grossman, Thomas A.
Publié: (2007)
par: Grossman, Thomas A.
Publié: (2007)
nvidia-pcm: A D-Bus-Driven Platform Configuration Manager for OpenBMC Environments
par: Singh, Harinder
Publié: (2026)
par: Singh, Harinder
Publié: (2026)
Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference
par: Dong, Wenxin, et autres
Publié: (2026)
par: Dong, Wenxin, et autres
Publié: (2026)
Synergy of Large Language Model and Model Driven Engineering for Automated Development of Centralized Vehicular Systems
par: Petrovic, Nenad, et autres
Publié: (2024)
par: Petrovic, Nenad, et autres
Publié: (2024)
Towards Single-System Illusion in Software-Defined Vehicles -- Automated, AI-Powered Workflow
par: Lebioda, Krzysztof, et autres
Publié: (2024)
par: Lebioda, Krzysztof, et autres
Publié: (2024)
LLM-Assisted Translation of Legacy FORTRAN Codes to C++: A Cross-Platform Study
par: Ranasinghe, Nishath Rajiv, et autres
Publié: (2025)
par: Ranasinghe, Nishath Rajiv, et autres
Publié: (2025)
Test-driven Software Experimentation with LASSO: an LLM Prompt Benchmarking Example
par: Kessel, Marcus
Publié: (2024)
par: Kessel, Marcus
Publié: (2024)
N-Version Assessment and Enhancement of Generative AI
par: Kessel, Marcus, et autres
Publié: (2024)
par: Kessel, Marcus, et autres
Publié: (2024)
Morescient GAI for Software Engineering (Extended Version)
par: Kessel, Marcus, et autres
Publié: (2024)
par: Kessel, Marcus, et autres
Publié: (2024)
Source Code Protection for Applications Written in Microsoft Excel and Google Spreadsheet
par: Grossman, Thomas A.
Publié: (2008)
par: Grossman, Thomas A.
Publié: (2008)
Binary-30K: A Heterogeneous Dataset for Deep Learning in Binary Analysis and Malware Detection
par: Bommarito II, Michael J.
Publié: (2025)
par: Bommarito II, Michael J.
Publié: (2025)
Documents similaires
-
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
par: Li, Xiangchen, et autres
Publié: (2026) -
SLED: A Speculative LLM Decoding Framework for Efficient Edge Serving
par: Li, Xiangchen, et autres
Publié: (2025) -
ICAN-Deploy: Identity-Stable Canary Deployment for Safety-Critical Embodied Agents
par: Qin, Xue, et autres
Publié: (2026) -
Secure coding for web applications: Frameworks, challenges, and the role of LLMs
par: Kiashemshaki, Kiana, et autres
Publié: (2025) -
Secure and Scalable Blockchain Voting: A Comparative Framework and the Role of Large Language Models
par: Kiashemshaki, Kiana, et autres
Publié: (2025)