SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
Fuente:
arXiv
Salvato in:
| Autori principali: | Wolfrath, Joel, Frink, Daniel, Chandra, Abhishek |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Biased Estimator for MinMax Sampling and Distributed Aggregation
di: Wolfrath, Joel, et al.
Pubblicazione: (2024)
di: Wolfrath, Joel, et al.
Pubblicazione: (2024)
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
di: Nguyen, Thanh-Tung, et al.
Pubblicazione: (2025)
di: Nguyen, Thanh-Tung, et al.
Pubblicazione: (2025)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
di: Cao, Jiahe, et al.
Pubblicazione: (2026)
di: Cao, Jiahe, et al.
Pubblicazione: (2026)
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
di: Pagonas, Nikos, et al.
Pubblicazione: (2025)
di: Pagonas, Nikos, et al.
Pubblicazione: (2025)
ACE-GNN: Adaptive GNN Co-Inference with System-Aware Scheduling in Dynamic Edge Environments
di: Zhou, Ao, et al.
Pubblicazione: (2025)
di: Zhou, Ao, et al.
Pubblicazione: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
di: Huang, Jinqi, et al.
Pubblicazione: (2025)
di: Huang, Jinqi, et al.
Pubblicazione: (2025)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
di: Li, Suyi, et al.
Pubblicazione: (2024)
di: Li, Suyi, et al.
Pubblicazione: (2024)
EcoServe: Designing Carbon-Aware AI Inference Systems
di: Li, Yueying, et al.
Pubblicazione: (2025)
di: Li, Yueying, et al.
Pubblicazione: (2025)
BrownoutServe: SLO-Aware Inference Serving under Bursty Workloads for MoE-based LLMs
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
di: Hu, Jianmin, et al.
Pubblicazione: (2025)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
di: Yuan, Yitao, et al.
Pubblicazione: (2025)
di: Yuan, Yitao, et al.
Pubblicazione: (2025)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
di: Peng, You, et al.
Pubblicazione: (2026)
di: Peng, You, et al.
Pubblicazione: (2026)
Preemption Aware Task Scheduling for Priority and Deadline Constrained DNN Inference Task Offloading in Homogeneous Mobile-Edge Networks
di: Cotter, Jamie, et al.
Pubblicazione: (2025)
di: Cotter, Jamie, et al.
Pubblicazione: (2025)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
di: Chow, Will
Pubblicazione: (2025)
di: Chow, Will
Pubblicazione: (2025)
Ding-Dong Ditch: Peeking Into Spot Instance Availability
di: Kim, Kyumin, et al.
Pubblicazione: (2026)
di: Kim, Kyumin, et al.
Pubblicazione: (2026)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
di: Zhan, Huiyou, et al.
Pubblicazione: (2025)
di: Zhan, Huiyou, et al.
Pubblicazione: (2025)
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
di: Raj, Suman, et al.
Pubblicazione: (2024)
di: Raj, Suman, et al.
Pubblicazione: (2024)
Modular Foundation Model Inference at the Edge: Network-Aware Microservice Optimization
di: Zhu, Juan, et al.
Pubblicazione: (2026)
di: Zhu, Juan, et al.
Pubblicazione: (2026)
Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge Computing
di: Sheng, Jinhao, et al.
Pubblicazione: (2025)
di: Sheng, Jinhao, et al.
Pubblicazione: (2025)
SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
di: Jaiswal, Shashwat, et al.
Pubblicazione: (2025)
DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling
di: Ahmad, Sohaib, et al.
Pubblicazione: (2024)
di: Ahmad, Sohaib, et al.
Pubblicazione: (2024)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
di: Rajashekar, Kolichala, et al.
Pubblicazione: (2025)
di: Rajashekar, Kolichala, et al.
Pubblicazione: (2025)
Environment-Aware Dynamic Pruning for Pipelined Edge Inference
di: O'Quinn, Austin, et al.
Pubblicazione: (2025)
di: O'Quinn, Austin, et al.
Pubblicazione: (2025)
TCM-Serve: Modality-aware Scheduling for Multimodal Large Language Model Inference
di: Papaioannou, Konstantinos, et al.
Pubblicazione: (2026)
di: Papaioannou, Konstantinos, et al.
Pubblicazione: (2026)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
di: Wang, Qipeng
Pubblicazione: (2026)
di: Wang, Qipeng
Pubblicazione: (2026)
Past-Future Scheduler for LLM Serving under SLA Guarantees
di: Gong, Ruihao, et al.
Pubblicazione: (2025)
di: Gong, Ruihao, et al.
Pubblicazione: (2025)
PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
di: Huang, Weizhe, et al.
Pubblicazione: (2025)
di: Huang, Weizhe, et al.
Pubblicazione: (2025)
FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
di: Wu, Li, et al.
Pubblicazione: (2025)
di: Wu, Li, et al.
Pubblicazione: (2025)
A Predictive and Synergistic Two-Layer Scheduling Framework for LLM Serving
di: Zhang, Yue, et al.
Pubblicazione: (2025)
di: Zhang, Yue, et al.
Pubblicazione: (2025)
Slice-Level Scheduling for High Throughput and Load Balanced LLM Serving
di: Cheng, Ke, et al.
Pubblicazione: (2024)
di: Cheng, Ke, et al.
Pubblicazione: (2024)
Cloud Native System for LLM Inference Serving
di: Xu, Minxian, et al.
Pubblicazione: (2025)
di: Xu, Minxian, et al.
Pubblicazione: (2025)
Serving Compound Inference Systems on Datacenter GPUs
di: Devata, Sriram, et al.
Pubblicazione: (2026)
di: Devata, Sriram, et al.
Pubblicazione: (2026)
Seer: Proactive Revenue-Aware Scheduling for Live Streaming Services in Crowdsourced Cloud-Edge Platforms
di: Huang, Shaoyuan, et al.
Pubblicazione: (2024)
di: Huang, Shaoyuan, et al.
Pubblicazione: (2024)
Learning to Schedule: A Supervised Learning Framework for Network-Aware Scheduling of Data-Intensive Workloads
di: Timilsina, Sankalpa, et al.
Pubblicazione: (2025)
di: Timilsina, Sankalpa, et al.
Pubblicazione: (2025)
GoodServe: Towards High-Goodput Serving of Agentic LLM Inferences over Heterogeneous Resources
di: Du, Boxiao, et al.
Pubblicazione: (2026)
di: Du, Boxiao, et al.
Pubblicazione: (2026)
FREESH: Fair, Resource- and Energy-Efficient Scheduling for LLM Serving on Heterogeneous GPUs
di: He, Xuan, et al.
Pubblicazione: (2025)
di: He, Xuan, et al.
Pubblicazione: (2025)
Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and Splitting
di: Zhao, Zhixin, et al.
Pubblicazione: (2024)
di: Zhao, Zhixin, et al.
Pubblicazione: (2024)
FineServe: Precision-Aware KV Slab and Two-Level Scheduling for Heterogeneous Precision LLM Serving
di: Bin, Kyungmin, et al.
Pubblicazione: (2025)
di: Bin, Kyungmin, et al.
Pubblicazione: (2025)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
di: Liu, Dong, et al.
Pubblicazione: (2025)
di: Liu, Dong, et al.
Pubblicazione: (2025)
Improved Decision Module Selection for Hierarchical Inference in Resource-Constrained Edge Devices
di: Behera, Adarsh Prasad, et al.
Pubblicazione: (2024)
di: Behera, Adarsh Prasad, et al.
Pubblicazione: (2024)
Documenti analoghi
-
A Biased Estimator for MinMax Sampling and Distributed Aggregation
di: Wolfrath, Joel, et al.
Pubblicazione: (2024) -
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
di: Nguyen, Thanh-Tung, et al.
Pubblicazione: (2025) -
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
di: Cao, Jiahe, et al.
Pubblicazione: (2026) -
Cortex: Workflow-Aware Resource Pooling and Scheduling for Agentic Serving
di: Pagonas, Nikos, et al.
Pubblicazione: (2025) -
ACE-GNN: Adaptive GNN Co-Inference with System-Aware Scheduling in Dynamic Edge Environments
di: Zhou, Ao, et al.
Pubblicazione: (2025)