Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Shibo, Goudarzi, Mohammad, Toosi, Adel Nadjaran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Resilience Evaluation of Kubernetes in Cloud-Edge Environments via Failure Injection
by: Chen, Zihao, et al.
Published: (2025)
by: Chen, Zihao, et al.
Published: (2025)
A Multi-Armed Bandit-Based Participant Selection Method for Federated Recommendation Systems
by: Liu, Jintao, et al.
Published: (2025)
by: Liu, Jintao, et al.
Published: (2025)
LLM-Driven Intent-Based Privacy-Aware Orchestration Across the Cloud-Edge Continuum
by: Su, Zijie, et al.
Published: (2026)
by: Su, Zijie, et al.
Published: (2026)
ReinFog: A Deep Reinforcement Learning Empowered Framework for Resource Management in Edge and Cloud Computing Environments
by: Wang, Zhiyu, et al.
Published: (2024)
by: Wang, Zhiyu, et al.
Published: (2024)
TF-DDRL: A Transformer-enhanced Distributed DRL Technique for Scheduling IoT Applications in Edge and Cloud Computing Environments
by: Wang, Zhiyu, et al.
Published: (2024)
by: Wang, Zhiyu, et al.
Published: (2024)
REACH: Reinforcement Learning for Adaptive Microservice Rescheduling in the Cloud-Edge Continuum
by: Bai, Xu, et al.
Published: (2025)
by: Bai, Xu, et al.
Published: (2025)
HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
by: Dong, Jiangwen, et al.
Published: (2025)
by: Dong, Jiangwen, et al.
Published: (2025)
GraphFlash: Enabling Fast and Elastic Graph Processing on Serverless Infrastructure
by: Zhao, Chen, et al.
Published: (2026)
by: Zhao, Chen, et al.
Published: (2026)
IntentContinuum: Using LLMs to Support Intent-Based Computing Across the Compute Continuum
by: Akbari, Negin, et al.
Published: (2025)
by: Akbari, Negin, et al.
Published: (2025)
TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
by: Wen, Linfeng, et al.
Published: (2024)
by: Wen, Linfeng, et al.
Published: (2024)
Personalizing Federated Learning for Hierarchical Edge Networks with Non-IID Data
by: Lee, Seunghyun, et al.
Published: (2025)
by: Lee, Seunghyun, et al.
Published: (2025)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026)
by: Kasnavieh, Hossein Hosseini, et al.
Published: (2026)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
by: Zhang, Yida, et al.
Published: (2026)
by: Zhang, Yida, et al.
Published: (2026)
Multi-Layer Scheduling for MoE-Based LLM Reasoning
by: Sun, Yifan, et al.
Published: (2026)
by: Sun, Yifan, et al.
Published: (2026)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
by: Yang, Zheming, et al.
Published: (2026)
by: Yang, Zheming, et al.
Published: (2026)
PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving
by: Bai, Xu, et al.
Published: (2026)
by: Bai, Xu, et al.
Published: (2026)
Performance and Security Aware Distributed Service Placement in Fog Computing
by: Goudarzi, Mohammad, et al.
Published: (2026)
by: Goudarzi, Mohammad, et al.
Published: (2026)
ECORE: Energy-Conscious Optimized Routing for Deep Learning Models at the Edge
by: Alqahtani, Daghash K., et al.
Published: (2025)
by: Alqahtani, Daghash K., et al.
Published: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
by: Wang, Qipeng
Published: (2026)
by: Wang, Qipeng
Published: (2026)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
by: Yang, Zheming, et al.
Published: (2025)
by: Yang, Zheming, et al.
Published: (2025)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
by: Tang, Lujie, et al.
Published: (2024)
by: Tang, Lujie, et al.
Published: (2024)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
by: Chow, Will
Published: (2025)
by: Chow, Will
Published: (2025)
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
by: Tayal, Mumuksh, et al.
Published: (2025)
by: Tayal, Mumuksh, et al.
Published: (2025)
PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
by: Han, Yunhe, et al.
Published: (2026)
by: Han, Yunhe, et al.
Published: (2026)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
by: Zhan, Huiyou, et al.
Published: (2025)
by: Zhan, Huiyou, et al.
Published: (2025)
Efficient Accelerated Graph Edit Distance Computation on GPU
by: Dabah, Adel, et al.
Published: (2026)
by: Dabah, Adel, et al.
Published: (2026)
Energy Metrics for Edge Microservice Request Placement Strategies
by: Toczé, Klervie, et al.
Published: (2025)
by: Toczé, Klervie, et al.
Published: (2025)
Multi-Objective Load Balancing for Heterogeneous Edge-Based Object Detection Systems
by: Alqahtani, Daghash K., et al.
Published: (2026)
by: Alqahtani, Daghash K., et al.
Published: (2026)
Energy-aware Distributed Microservice Request Placement at the Edge
by: Toczé, Klervie, et al.
Published: (2024)
by: Toczé, Klervie, et al.
Published: (2024)
ML-ECS: A Collaborative Multimodal Learning Framework for Edge-Cloud Synergies
by: Liu, Yuze, et al.
Published: (2026)
by: Liu, Yuze, et al.
Published: (2026)
Propius: A Platform for Collaborative Machine Learning across the Edge and the Cloud
by: Ding, Eric
Published: (2025)
by: Ding, Eric
Published: (2025)
Cloud Native System for LLM Inference Serving
by: Xu, Minxian, et al.
Published: (2025)
by: Xu, Minxian, et al.
Published: (2025)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
by: Zhang, Zhexiang, et al.
Published: (2025)
by: Zhang, Zhexiang, et al.
Published: (2025)
Towards Seamless Serverless Computing Across an Edge-Cloud Continuum
by: Simion, Emilian, et al.
Published: (2024)
by: Simion, Emilian, et al.
Published: (2024)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
by: Arya, Mayank, et al.
Published: (2025)
by: Arya, Mayank, et al.
Published: (2025)
Toward Sustainability-Aware LLM Inference on Edge Clusters
by: Rajashekar, Kolichala, et al.
Published: (2025)
by: Rajashekar, Kolichala, et al.
Published: (2025)
A Knowledge Distillation-empowered Adaptive Federated Reinforcement Learning Framework for Multi-Domain IoT Applications Scheduling
by: Wang, Zhiyu, et al.
Published: (2025)
by: Wang, Zhiyu, et al.
Published: (2025)
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
by: Raj, Suman, et al.
Published: (2024)
by: Raj, Suman, et al.
Published: (2024)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
by: Sun, Mingyu, et al.
Published: (2025)
by: Sun, Mingyu, et al.
Published: (2025)
Similar Items
-
Resilience Evaluation of Kubernetes in Cloud-Edge Environments via Failure Injection
by: Chen, Zihao, et al.
Published: (2025) -
A Multi-Armed Bandit-Based Participant Selection Method for Federated Recommendation Systems
by: Liu, Jintao, et al.
Published: (2025) -
LLM-Driven Intent-Based Privacy-Aware Orchestration Across the Cloud-Edge Continuum
by: Su, Zijie, et al.
Published: (2026) -
ReinFog: A Deep Reinforcement Learning Empowered Framework for Resource Management in Edge and Cloud Computing Environments
by: Wang, Zhiyu, et al.
Published: (2024) -
TF-DDRL: A Transformer-enhanced Distributed DRL Technique for Scheduling IoT Applications in Edge and Cloud Computing Environments
by: Wang, Zhiyu, et al.
Published: (2024)