Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Rui, Ouyang, Tao, Zeng, Liekang, Liao, Guocheng, Zhou, Zhi, Chen, Xu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
von: Hong, Guihang, et al.
Veröffentlicht: (2025)
von: Hong, Guihang, et al.
Veröffentlicht: (2025)
Adaptive Device-Edge Collaboration on DNN Inference in AIoT: A Digital Twin-Assisted Approach
von: Hu, Shisheng, et al.
Veröffentlicht: (2024)
von: Hu, Shisheng, et al.
Veröffentlicht: (2024)
Implementation of Big AI Models for Wireless Networks with Collaborative Edge Computing
von: Zeng, Liekang, et al.
Veröffentlicht: (2024)
von: Zeng, Liekang, et al.
Veröffentlicht: (2024)
Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
von: Ye, Shengyuan, et al.
Veröffentlicht: (2025)
von: Ye, Shengyuan, et al.
Veröffentlicht: (2025)
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
von: Ye, Shengyuan, et al.
Veröffentlicht: (2025)
von: Ye, Shengyuan, et al.
Veröffentlicht: (2025)
HarmonyBatch: Batching multi-SLO DNN Inference with Heterogeneous Serverless Functions
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
von: Chen, Jiabin, et al.
Veröffentlicht: (2024)
Learning the Optimal Path and DNN Partition for Collaborative Edge Inference
von: Huang, Yin, et al.
Veröffentlicht: (2024)
von: Huang, Yin, et al.
Veröffentlicht: (2024)
A Survey on Collaborative DNN Inference for Edge Intelligence
von: Ren, Weiqing, et al.
Veröffentlicht: (2022)
von: Ren, Weiqing, et al.
Veröffentlicht: (2022)
Infer-EDGE: Dynamic DNN Inference Optimization in 'Just-in-time' Edge-AI Implementations
von: Mounesan, Motahare, et al.
Veröffentlicht: (2025)
von: Mounesan, Motahare, et al.
Veröffentlicht: (2025)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
von: Cao, Jiahe, et al.
Veröffentlicht: (2026)
von: Cao, Jiahe, et al.
Veröffentlicht: (2026)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)
von: Zhang, Mingjin, et al.
Veröffentlicht: (2024)
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2023)
von: K., Prashanthi S., et al.
Veröffentlicht: (2023)
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2025)
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2025)
Collaborative Inference in DNN-based Satellite Systems with Dynamic Task Streams
von: Guan, Jinglong, et al.
Veröffentlicht: (2023)
von: Guan, Jinglong, et al.
Veröffentlicht: (2023)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
von: Chen, Aodong, et al.
Veröffentlicht: (2023)
von: Chen, Aodong, et al.
Veröffentlicht: (2023)
Asteroid: Resource-Efficient Hybrid Pipeline Parallelism for Collaborative DNN Training on Heterogeneous Edge Devices
von: Ye, Shengyuan, et al.
Veröffentlicht: (2024)
von: Ye, Shengyuan, et al.
Veröffentlicht: (2024)
Adaptive Heuristics for Scheduling DNN Inferencing on Edge and Cloud for Personalized UAV Fleets
von: Raj, Suman, et al.
Veröffentlicht: (2024)
von: Raj, Suman, et al.
Veröffentlicht: (2024)
Where to Split? A Pareto-Front Analysis of DNN Partitioning for Edge Inference
von: Masud, Adiba, et al.
Veröffentlicht: (2026)
von: Masud, Adiba, et al.
Veröffentlicht: (2026)
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
von: Liu, Kaiwei, et al.
Veröffentlicht: (2025)
von: Liu, Kaiwei, et al.
Veröffentlicht: (2025)
Collaborative Satellite Computing through Adaptive DNN Task Splitting and Offloading
von: Peng, Shifeng, et al.
Veröffentlicht: (2024)
von: Peng, Shifeng, et al.
Veröffentlicht: (2024)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
von: Zhang, Yida, et al.
Veröffentlicht: (2026)
Preemption Aware Task Scheduling for Priority and Deadline Constrained DNN Inference Task Offloading in Homogeneous Mobile-Edge Networks
von: Cotter, Jamie, et al.
Veröffentlicht: (2025)
von: Cotter, Jamie, et al.
Veröffentlicht: (2025)
CCRSat: A Collaborative Computation Reuse Framework for Satellite Edge Computing Networks
von: Zhang, Ye, et al.
Veröffentlicht: (2025)
von: Zhang, Ye, et al.
Veröffentlicht: (2025)
Resource-Efficient Personal Large Language Models Fine-Tuning with Collaborative Edge Computing
von: Ye, Shengyuan, et al.
Veröffentlicht: (2024)
von: Ye, Shengyuan, et al.
Veröffentlicht: (2024)
Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
von: Li, Senyao, et al.
Veröffentlicht: (2025)
von: Li, Senyao, et al.
Veröffentlicht: (2025)
AdaBridge: Dynamic Data and Computation Reuse for Efficient Multi-task DNN Co-evolution in Edge Systems
von: Wang, Lehao, et al.
Veröffentlicht: (2024)
von: Wang, Lehao, et al.
Veröffentlicht: (2024)
Nezha: Breaking Multi-Rail Network Barriers for Distributed DNN Training
von: Yu, Enda, et al.
Veröffentlicht: (2024)
von: Yu, Enda, et al.
Veröffentlicht: (2024)
Collaborative Inference for Large Models with Task Offloading and Early Exiting
von: Xie, Zuan, et al.
Veröffentlicht: (2024)
von: Xie, Zuan, et al.
Veröffentlicht: (2024)
Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference
von: Ye, Shengyuan, et al.
Veröffentlicht: (2024)
von: Ye, Shengyuan, et al.
Veröffentlicht: (2024)
Modular Foundation Model Inference at the Edge: Network-Aware Microservice Optimization
von: Zhu, Juan, et al.
Veröffentlicht: (2026)
von: Zhu, Juan, et al.
Veröffentlicht: (2026)
Parallel Collaborative ADMM Privacy Computing and Adaptive GPU Acceleration for Distributed Edge Networks
von: Xia, Mengchun, et al.
Veröffentlicht: (2026)
von: Xia, Mengchun, et al.
Veröffentlicht: (2026)
Barycentric Coded Distributed Computing with Flexible Recovery Threshold for Collaborative Mobile Edge Computing
von: Qiu, Houming, et al.
Veröffentlicht: (2025)
von: Qiu, Houming, et al.
Veröffentlicht: (2025)
ECC-SNN: Cost-Effective Edge-Cloud Collaboration for Spiking Neural Networks
von: Yu, Di, et al.
Veröffentlicht: (2025)
von: Yu, Di, et al.
Veröffentlicht: (2025)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
von: Sun, Mingyu, et al.
Veröffentlicht: (2025)
von: Sun, Mingyu, et al.
Veröffentlicht: (2025)
Collaborative Resource Management and Workloads Scheduling in Cloud-Assisted Mobile Edge Computing across Timescales
von: Tang, Lujie, et al.
Veröffentlicht: (2024)
von: Tang, Lujie, et al.
Veröffentlicht: (2024)
AdaOper: Energy-efficient and Responsive Concurrent DNN Inference on Mobile Devices
von: Lin, Zheng, et al.
Veröffentlicht: (2024)
von: Lin, Zheng, et al.
Veröffentlicht: (2024)
Pagoda: An Energy and Time Roofline Study for DNN Workloads on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)
Many Hands Make Light Work: Accelerating Edge Inference via Multi-Client Collaborative Caching
von: Liang, Wenyi, et al.
Veröffentlicht: (2024)
von: Liang, Wenyi, et al.
Veröffentlicht: (2024)
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
von: Hong, Guihang, et al.
Veröffentlicht: (2025) -
Adaptive Device-Edge Collaboration on DNN Inference in AIoT: A Digital Twin-Assisted Approach
von: Hu, Shisheng, et al.
Veröffentlicht: (2024) -
Implementation of Big AI Models for Wireless Networks with Collaborative Edge Computing
von: Zeng, Liekang, et al.
Veröffentlicht: (2024) -
Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
von: Ye, Shengyuan, et al.
Veröffentlicht: (2025) -
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
von: K., Prashanthi S., et al.
Veröffentlicht: (2025)