Priority-Aware Model-Distributed Inference at Edge Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Teng, Seferoglu, Hulya |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Early-Exit meets Model-Distributed Inference at Edge Networks
by: Colocrese, Marco, et al.
Published: (2024)
by: Colocrese, Marco, et al.
Published: (2024)
DIGEST: Fast and Communication Efficient Decentralized Learning with Local Updates
by: Gholami, Peyman, et al.
Published: (2023)
by: Gholami, Peyman, et al.
Published: (2023)
Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference
by: Siavashi, Mohammad, et al.
Published: (2025)
by: Siavashi, Mohammad, et al.
Published: (2025)
Intelligent Orchestration of Distributed Large Foundation Model Inference at the Edge
by: Koch, Fernando, et al.
Published: (2025)
by: Koch, Fernando, et al.
Published: (2025)
ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer Inference
by: Meng, Han, et al.
Published: (2026)
by: Meng, Han, et al.
Published: (2026)
Distributed Convolutional Neural Network Training on Mobile and Edge Clusters
by: Rama, Pranav, et al.
Published: (2024)
by: Rama, Pranav, et al.
Published: (2024)
Robust Federated Learning against Model Perturbation in Edge Networks
by: Jin, Dongzi, et al.
Published: (2025)
by: Jin, Dongzi, et al.
Published: (2025)
DistrEE: Distributed Early Exit of Deep Neural Network Inference on Edge Devices
by: Peng, Xian, et al.
Published: (2025)
by: Peng, Xian, et al.
Published: (2025)
Preemption Aware Task Scheduling for Priority and Deadline Constrained DNN Inference Task Offloading in Homogeneous Mobile-Edge Networks
by: Cotter, Jamie, et al.
Published: (2025)
by: Cotter, Jamie, et al.
Published: (2025)
Fast Distributed Inference Serving for Large Language Models
by: Wu, Bingyang, et al.
Published: (2023)
by: Wu, Bingyang, et al.
Published: (2023)
A Survey on Collaborative DNN Inference for Edge Intelligence
by: Ren, Weiqing, et al.
Published: (2022)
by: Ren, Weiqing, et al.
Published: (2022)
DSD: A Distributed Speculative Decoding Solution for Edge-Cloud Agile Large Model Serving
by: Yu, Fengze, et al.
Published: (2025)
by: Yu, Fengze, et al.
Published: (2025)
Adaptive Stream Processing on Edge Devices through Active Inference
by: Sedlak, Boris, et al.
Published: (2024)
by: Sedlak, Boris, et al.
Published: (2024)
Learning the Optimal Path and DNN Partition for Collaborative Edge Inference
by: Huang, Yin, et al.
Published: (2024)
by: Huang, Yin, et al.
Published: (2024)
NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning
by: Wang, Irene, et al.
Published: (2026)
by: Wang, Irene, et al.
Published: (2026)
Federated Attention: A Distributed Paradigm for Collaborative LLM Inference over Edge Networks
by: Deng, Xiumei, et al.
Published: (2025)
by: Deng, Xiumei, et al.
Published: (2025)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
by: Liu, Mengfan, et al.
Published: (2025)
by: Liu, Mengfan, et al.
Published: (2025)
Model Partition and Resource Allocation for Split Learning in Vehicular Edge Networks
by: Yu, Lu, et al.
Published: (2024)
by: Yu, Lu, et al.
Published: (2024)
Multi-DNN Inference of Sparse Models on Edge SoCs
by: Luo, Jiawei, et al.
Published: (2026)
by: Luo, Jiawei, et al.
Published: (2026)
Leveraging Foundation Models for Efficient Federated Learning in Resource-restricted Edge Networks
by: Atapour, S. Kawa, et al.
Published: (2024)
by: Atapour, S. Kawa, et al.
Published: (2024)
TSFLora: Token-Compressed Split Fine-Tuning for Wireless Edge Networks
by: Qiang, Xianke, et al.
Published: (2026)
by: Qiang, Xianke, et al.
Published: (2026)
Towards Integrated Fine-tuning and Inference when Generative AI meets Edge Intelligence
by: Chen, Ning, et al.
Published: (2024)
by: Chen, Ning, et al.
Published: (2024)
MobiZO: Enabling Efficient LLM Fine-Tuning at the Edge via Inference Engines
by: Gao, Lei, et al.
Published: (2024)
by: Gao, Lei, et al.
Published: (2024)
Heterogeneity-Aware Cooperative Federated Edge Learning with Adaptive Computation and Communication Compression
by: Zhang, Zhenxiao, et al.
Published: (2024)
by: Zhang, Zhenxiao, et al.
Published: (2024)
MultiTASC++: A Continuously Adaptive Scheduler for Edge-Based Multi-Device Cascade Inference
by: Nikolaidis, Sokratis, et al.
Published: (2024)
by: Nikolaidis, Sokratis, et al.
Published: (2024)
ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference
by: Oh, Hyungjun, et al.
Published: (2024)
by: Oh, Hyungjun, et al.
Published: (2024)
RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
Lightweight Federated Learning over Wireless Edge Networks
by: Hou, Xiangwang, et al.
Published: (2025)
by: Hou, Xiangwang, et al.
Published: (2025)
Modular Foundation Model Inference at the Edge: Network-Aware Microservice Optimization
by: Zhu, Juan, et al.
Published: (2026)
by: Zhu, Juan, et al.
Published: (2026)
Split CNN Inference on Networked Microcontrollers
by: Lu, Junyu, et al.
Published: (2026)
by: Lu, Junyu, et al.
Published: (2026)
FedTeddi: Temporal Drift and Divergence Aware Scheduling for Timely Federated Edge Learning
by: Bai, Yuxuan, et al.
Published: (2025)
by: Bai, Yuxuan, et al.
Published: (2025)
Deal: Distributed End-to-End GNN Inference for All Nodes
by: Chen, Shiyang, et al.
Published: (2025)
by: Chen, Shiyang, et al.
Published: (2025)
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
by: Tummalapalli, Pranay, et al.
Published: (2026)
by: Tummalapalli, Pranay, et al.
Published: (2026)
Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms
by: Xu, Ao, et al.
Published: (2025)
by: Xu, Ao, et al.
Published: (2025)
DALI: A Workload-Aware Offloading Framework for Efficient MoE Inference on Local PCs
by: Zhu, Zeyu, et al.
Published: (2026)
by: Zhu, Zeyu, et al.
Published: (2026)
Optimizing Performance on Trinity Utilizing Machine Learning, Proxy Applications and Scheduling Priorities
by: Romero, Phil
Published: (2024)
by: Romero, Phil
Published: (2024)
Accelerating Local LLMs on Resource-Constrained Edge Devices via Distributed Prompt Caching
by: Matsutani, Hiroki, et al.
Published: (2026)
by: Matsutani, Hiroki, et al.
Published: (2026)
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
by: Gond, Raja, et al.
Published: (2025)
by: Gond, Raja, et al.
Published: (2025)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
by: Xu, Guanyu, et al.
Published: (2025)
by: Xu, Guanyu, et al.
Published: (2025)
DisCEdge: Distributed Context Management for Large Language Models at the Edge
by: Malekabbasi, Mohammadreza, et al.
Published: (2025)
by: Malekabbasi, Mohammadreza, et al.
Published: (2025)
Similar Items
-
Early-Exit meets Model-Distributed Inference at Edge Networks
by: Colocrese, Marco, et al.
Published: (2024) -
DIGEST: Fast and Communication Efficient Decentralized Learning with Local Updates
by: Gholami, Peyman, et al.
Published: (2023) -
Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference
by: Siavashi, Mohammad, et al.
Published: (2025) -
Intelligent Orchestration of Distributed Large Foundation Model Inference at the Edge
by: Koch, Fernando, et al.
Published: (2025) -
ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer Inference
by: Meng, Han, et al.
Published: (2026)