Saved in:
| Main Authors: | Rajashekar, Kolichala, Sharghivand, Nafiseh, Prodan, Radu, Farahani, Reza |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.04088 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ClusterLess: Deadline-Aware Serverless Workflow Orchestration on Federated Edge Clusters
by: Farahani, Reza, et al.
Published: (2026)
by: Farahani, Reza, et al.
Published: (2026)
Serverless Everywhere: A Comparative Analysis of WebAssembly Workflows Across Browser, Edge, and Cloud
by: Colosi, Mario, et al.
Published: (2025)
by: Colosi, Mario, et al.
Published: (2025)
Osmotic Learning: A Self-Supervised Paradigm for Decentralized Contextual Data Representation
by: Colosi, Mario, et al.
Published: (2025)
by: Colosi, Mario, et al.
Published: (2025)
SCAREY: Location-Aware Service Lifecycle Management
by: Horvath, Kurt, et al.
Published: (2025)
by: Horvath, Kurt, et al.
Published: (2025)
Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
by: Nicolae, Radu, et al.
Published: (2026)
by: Nicolae, Radu, et al.
Published: (2026)
ADApt: Edge Device Anomaly Detection and Microservice Replica Prediction
by: Mehran, Narges, et al.
Published: (2025)
by: Mehran, Narges, et al.
Published: (2025)
Federated Distillation on Edge Devices: Efficient Client-Side Filtering for Non-IID Data
by: Mujtaba, Ahmed, et al.
Published: (2025)
by: Mujtaba, Ahmed, et al.
Published: (2025)
Orchestrating Serverless Applications in the Edge Cloud Space Continuum: What Breaks and What is Next?
by: Malazi, Hadi Tabatabaee, et al.
Published: (2026)
by: Malazi, Hadi Tabatabaee, et al.
Published: (2026)
Scale: Deep Reinforcement Learning for Container Scheduling in Serverless Edge Computing
by: Chen, Chen, et al.
Published: (2026)
by: Chen, Chen, et al.
Published: (2026)
Enhancing Traffic Safety with AI and 6G: Latency Requirements and Real-Time Threat Detection
by: Horvath, Kurt, et al.
Published: (2025)
by: Horvath, Kurt, et al.
Published: (2025)
FedADAS: Communication-Efficient Federated Distillation for On-Device Driver Yawn Recognition in Vehicular Networks
by: Mujtaba, Ahmed, et al.
Published: (2026)
by: Mujtaba, Ahmed, et al.
Published: (2026)
DEEP: Edge-based Dataflow Processing with Hybrid Docker Hub and Regional Registries
by: Mehran, Narges, et al.
Published: (2025)
by: Mehran, Narges, et al.
Published: (2025)
Scheduling of Distributed Applications on the Computing Continuum: A Survey
by: Mehran, Narges, et al.
Published: (2024)
by: Mehran, Narges, et al.
Published: (2024)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
by: Yang, Zheming, et al.
Published: (2026)
by: Yang, Zheming, et al.
Published: (2026)
Environment-Aware Dynamic Pruning for Pipelined Edge Inference
by: O'Quinn, Austin, et al.
Published: (2025)
by: O'Quinn, Austin, et al.
Published: (2025)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
by: Arya, Mayank, et al.
Published: (2025)
by: Arya, Mayank, et al.
Published: (2025)
OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics
by: Nguyen, Thanh-Tung, et al.
Published: (2025)
by: Nguyen, Thanh-Tung, et al.
Published: (2025)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
by: Khoshsirat, Aria, et al.
Published: (2024)
by: Khoshsirat, Aria, et al.
Published: (2024)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
by: Yang, Zheming, et al.
Published: (2025)
by: Yang, Zheming, et al.
Published: (2025)
Modular Foundation Model Inference at the Edge: Network-Aware Microservice Optimization
by: Zhu, Juan, et al.
Published: (2026)
by: Zhu, Juan, et al.
Published: (2026)
SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
by: Chow, Will
Published: (2025)
by: Chow, Will
Published: (2025)
SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge
by: Wolfrath, Joel, et al.
Published: (2025)
by: Wolfrath, Joel, et al.
Published: (2025)
Argus: Token Aware Distributed LLM Inference Optimization
by: Wu, Panlong, et al.
Published: (2025)
by: Wu, Panlong, et al.
Published: (2025)
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism
by: Wei, Jinhui, et al.
Published: (2025)
by: Wei, Jinhui, et al.
Published: (2025)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
by: Sun, Mingyu, et al.
Published: (2025)
by: Sun, Mingyu, et al.
Published: (2025)
Efficient Routing of Inference Requests across LLM Instances in Cloud-Edge Computing
by: Yu, Shibo, et al.
Published: (2025)
by: Yu, Shibo, et al.
Published: (2025)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
by: Xu, Chuhao, et al.
Published: (2025)
by: Xu, Chuhao, et al.
Published: (2025)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
by: Dong, Jiangwen, et al.
Published: (2025)
by: Dong, Jiangwen, et al.
Published: (2025)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
by: Zhang, Yida, et al.
Published: (2026)
by: Zhang, Yida, et al.
Published: (2026)
ACE-GNN: Adaptive GNN Co-Inference with System-Aware Scheduling in Dynamic Edge Environments
by: Zhou, Ao, et al.
Published: (2025)
by: Zhou, Ao, et al.
Published: (2025)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
by: Zhan, Huiyou, et al.
Published: (2025)
by: Zhan, Huiyou, et al.
Published: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
by: Wang, Qipeng
Published: (2026)
by: Wang, Qipeng
Published: (2026)
LLM-Driven Intent-Based Privacy-Aware Orchestration Across the Cloud-Edge Continuum
by: Su, Zijie, et al.
Published: (2026)
by: Su, Zijie, et al.
Published: (2026)
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
by: Tummalapalli, Pranay, et al.
Published: (2026)
by: Tummalapalli, Pranay, et al.
Published: (2026)
BandPilot: Towards Performance- and Contention-Aware GPU Dispatching in AI Clusters
by: Zhang, Kunming, et al.
Published: (2025)
by: Zhang, Kunming, et al.
Published: (2025)
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
by: Argerich, Mauricio Fadel, et al.
Published: (2026)
by: Argerich, Mauricio Fadel, et al.
Published: (2026)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
by: Cao, Jiahe, et al.
Published: (2026)
by: Cao, Jiahe, et al.
Published: (2026)
Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
by: Chang, Zihan, et al.
Published: (2024)
by: Chang, Zihan, et al.
Published: (2024)
Similar Items
-
ClusterLess: Deadline-Aware Serverless Workflow Orchestration on Federated Edge Clusters
by: Farahani, Reza, et al.
Published: (2026) -
Serverless Everywhere: A Comparative Analysis of WebAssembly Workflows Across Browser, Edge, and Cloud
by: Colosi, Mario, et al.
Published: (2025) -
Osmotic Learning: A Self-Supervised Paradigm for Decentralized Contextual Data Representation
by: Colosi, Mario, et al.
Published: (2025) -
SCAREY: Location-Aware Service Lifecycle Management
by: Horvath, Kurt, et al.
Published: (2025) -
Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
by: Nicolae, Radu, et al.
Published: (2026)