PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Zheming, Yang, Yuanhao, Zhao, Chang, Guo, Qi, He, Wenkai, Ji, Wen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-Source Coflow Scheduling in Collaborative Edge Computing with Multihop Network
por: Sahni, Yuvraj, et al.
Publicado: (2024)
por: Sahni, Yuvraj, et al.
Publicado: (2024)
Recursive Offloading for LLM Serving in Multi-tier Networks
por: Wu, Zhiyuan, et al.
Publicado: (2025)
por: Wu, Zhiyuan, et al.
Publicado: (2025)
Multitier Service Migration Framework Based on Mobility Prediction in Mobile Edge Computing
por: Yang, Run, et al.
Publicado: (2024)
por: Yang, Run, et al.
Publicado: (2024)
Multi-stage Flow Scheduling for LLM Serving
por: Sun, Yijun, et al.
Publicado: (2026)
por: Sun, Yijun, et al.
Publicado: (2026)
EdgeTimer: Adaptive Multi-Timescale Scheduling in Mobile Edge Computing with Deep Reinforcement Learning
por: Hao, Yijun, et al.
Publicado: (2024)
por: Hao, Yijun, et al.
Publicado: (2024)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
por: Yang, Zheming, et al.
Publicado: (2025)
por: Yang, Zheming, et al.
Publicado: (2025)
Contention-Aware Microservice Deployment in Collaborative Mobile Edge Networks
por: Ge, Xinlei, et al.
Publicado: (2024)
por: Ge, Xinlei, et al.
Publicado: (2024)
Placing Timely Refreshing Services at the Network Edge
por: Li, Xishuo, et al.
Publicado: (2024)
por: Li, Xishuo, et al.
Publicado: (2024)
TriCloudEdge: A multi-layer Cloud Continuum
por: Violettas, George, et al.
Publicado: (2026)
por: Violettas, George, et al.
Publicado: (2026)
R2E-VID: Two-Stage Robust Routing via Temporal Gating for Elastic Edge-Cloud Video Inference
por: Yang, Zheming, et al.
Publicado: (2026)
por: Yang, Zheming, et al.
Publicado: (2026)
Carbon-Aware Temporal Data Transfer Scheduling Across Cloud Datacenters
por: Rodrigues, Elvis, et al.
Publicado: (2025)
por: Rodrigues, Elvis, et al.
Publicado: (2025)
An Open-Source Experimentation Framework for the Edge Cloud Continuum
por: Koukis, Georgios, et al.
Publicado: (2024)
por: Koukis, Georgios, et al.
Publicado: (2024)
HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge Network
por: Zheng, Peirong, et al.
Publicado: (2026)
por: Zheng, Peirong, et al.
Publicado: (2026)
SplitLLM: Collaborative Inference of LLMs for Model Placement and Throughput Optimization
por: Mudvari, Akrit, et al.
Publicado: (2024)
por: Mudvari, Akrit, et al.
Publicado: (2024)
Dynamic DAG-Application Scheduling for Multi-Tier Edge Computing in Heterogeneous Networks
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
An Online Fragmentation-Aware GPU Scheduler for Multi-Tenant MIG-based Clouds
por: Zambianco, Marco, et al.
Publicado: (2025)
por: Zambianco, Marco, et al.
Publicado: (2025)
Temporal-Aware GPU Resource Allocation for Distributed LLM Inference via Reinforcement Learning
por: Du, Chengze, et al.
Publicado: (2025)
por: Du, Chengze, et al.
Publicado: (2025)
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
por: Luo, Haoxiang, et al.
Publicado: (2025)
por: Luo, Haoxiang, et al.
Publicado: (2025)
Towards Timely Video Analytics Services at the Network Edge
por: Li, Xishuo, et al.
Publicado: (2024)
por: Li, Xishuo, et al.
Publicado: (2024)
Meili: Enabling SmartNIC as a Service in the Cloud
por: Su, Qiang, et al.
Publicado: (2023)
por: Su, Qiang, et al.
Publicado: (2023)
ClusterSlice: A Zero-touch Deployment Platform for the Edge Cloud Continuum
por: Mamatas, Lefteris, et al.
Publicado: (2024)
por: Mamatas, Lefteris, et al.
Publicado: (2024)
PeerSync: Accelerating Containerized Service Delivery at the Network Edge
por: Deng, Yinuo, et al.
Publicado: (2025)
por: Deng, Yinuo, et al.
Publicado: (2025)
A Task Decomposition and Planning Framework for Efficient LLM Inference in AI-Enabled WiFi-Offload Networks
por: Han, Mingqi, et al.
Publicado: (2026)
por: Han, Mingqi, et al.
Publicado: (2026)
Optimal Multi-Constrained Workflow Scheduling for Cyber-Physical Systems in the Edge-Cloud Continuum
por: Kouloumpris, Andreas, et al.
Publicado: (2025)
por: Kouloumpris, Andreas, et al.
Publicado: (2025)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
por: Yang, Zheming, et al.
Publicado: (2026)
por: Yang, Zheming, et al.
Publicado: (2026)
Causal Inference for Quantifying Noisy Neighbor Effects in Multi-Tenant Cloud Environments
por: Schiavo, Philipe S., et al.
Publicado: (2026)
por: Schiavo, Philipe S., et al.
Publicado: (2026)
A Study on 5G Network Slice Isolation Based on Native Cloud and Edge Computing Tools
por: Andrade, Maiko, et al.
Publicado: (2025)
por: Andrade, Maiko, et al.
Publicado: (2025)
Quality-of-Service Aware LLM Routing for Edge Computing with Multiple Experts
por: Yang, Jin, et al.
Publicado: (2025)
por: Yang, Jin, et al.
Publicado: (2025)
Hurry: Dynamic Collaborative Framework For Low-orbit Mega-Constellation Data Downloading
por: Luo, Handong, et al.
Publicado: (2024)
por: Luo, Handong, et al.
Publicado: (2024)
SPARC-LoRa: A Scalable, Power-efficient, Affordable, Reliable, and Cloud Service-enabled LoRa Networking System for Agriculture Applications
por: Wang, Xi, et al.
Publicado: (2024)
por: Wang, Xi, et al.
Publicado: (2024)
Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
por: Ye, Shengyuan, et al.
Publicado: (2025)
por: Ye, Shengyuan, et al.
Publicado: (2025)
Topology-aware Microservice Architecture in Edge Networks: Deployment Optimization and Implementation
por: Chen, Yuang, et al.
Publicado: (2025)
por: Chen, Yuang, et al.
Publicado: (2025)
MultiPath Memory Access: Breaking Host-GPU Bandwidth Bottlenecks in LLM Services
por: Tang, Lingfeng, et al.
Publicado: (2025)
por: Tang, Lingfeng, et al.
Publicado: (2025)
Arcturus: A Cloud Overlay Network for Global Accelerator with Enhanced Performance and Stability
por: Liu, Matthew Yang, et al.
Publicado: (2025)
por: Liu, Matthew Yang, et al.
Publicado: (2025)
CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure
por: Ding, Eric, et al.
Publicado: (2026)
por: Ding, Eric, et al.
Publicado: (2026)
Urgent Edge Computing
por: Dazzi, Patrizio, et al.
Publicado: (2024)
por: Dazzi, Patrizio, et al.
Publicado: (2024)
1.5 Million Messages Per Second on 3 Machines: Benchmarking and Latency Optimization of Apache Pulsar at Enterprise Scale
por: Mukkolakkal, Muhamed Ramees Cheriya
Publicado: (2026)
por: Mukkolakkal, Muhamed Ramees Cheriya
Publicado: (2026)
AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
por: Wei, Yufan, et al.
Publicado: (2025)
por: Wei, Yufan, et al.
Publicado: (2025)
FAST: An Efficient Scheduler for All-to-All GPU Communication
por: Lei, Yiran, et al.
Publicado: (2025)
por: Lei, Yiran, et al.
Publicado: (2025)
Edge Offloading in Smart Grid
por: Arcas, Gabriel Ioan, et al.
Publicado: (2024)
por: Arcas, Gabriel Ioan, et al.
Publicado: (2024)
Ejemplares similares
-
Multi-Source Coflow Scheduling in Collaborative Edge Computing with Multihop Network
por: Sahni, Yuvraj, et al.
Publicado: (2024) -
Recursive Offloading for LLM Serving in Multi-tier Networks
por: Wu, Zhiyuan, et al.
Publicado: (2025) -
Multitier Service Migration Framework Based on Mobility Prediction in Mobile Edge Computing
por: Yang, Run, et al.
Publicado: (2024) -
Multi-stage Flow Scheduling for LLM Serving
por: Sun, Yijun, et al.
Publicado: (2026) -
EdgeTimer: Adaptive Multi-Timescale Scheduling in Mobile Edge Computing with Deep Reinforcement Learning
por: Hao, Yijun, et al.
Publicado: (2024)