Adaptive Workload Distribution for Accuracy-aware DNN Inference on Collaborative Edge Platforms
Fuente:
arXiv
Saved in:
| Main Authors: | Taufique, Zain, Miele, Antonio, Liljeberg, Pasi, Kanduri, Anil |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge Platforms
by: Taufique, Zain, et al.
Published: (2024)
by: Taufique, Zain, et al.
Published: (2024)
Twill: Scheduling Compound AI Systems on Heterogeneous Mobile Edge Platforms
by: Taufique, Zain, et al.
Published: (2025)
by: Taufique, Zain, et al.
Published: (2025)
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
by: Li, Xiangchen, et al.
Published: (2025)
by: Li, Xiangchen, et al.
Published: (2025)
Energy-Optimized Scheduling for AIoT Workloads Using TOPSIS
by: Pradeep, Preethika, et al.
Published: (2025)
by: Pradeep, Preethika, et al.
Published: (2025)
Workload composition smooths aggregate power demand while sustaining short-horizon ramps in AI data centers
by: Majumder, Subir, et al.
Published: (2026)
by: Majumder, Subir, et al.
Published: (2026)
AARC: Automated Affinity-aware Resource Configuration for Serverless Workflows
by: Jin, Lingxiao, et al.
Published: (2025)
by: Jin, Lingxiao, et al.
Published: (2025)
AdapTBF: Decentralized Bandwidth Control via Adaptive Token Borrowing for HPC Storage
by: Rashid, Md Hasanur, et al.
Published: (2026)
by: Rashid, Md Hasanur, et al.
Published: (2026)
Characterizing Accuracy Trade-offs of EEG Applications on Embedded HMPs
by: Taufique, Zain, et al.
Published: (2024)
by: Taufique, Zain, et al.
Published: (2024)
Turning AI Data Centers into Grid-Interactive Assets: Results from a Field Demonstration in Phoenix, Arizona
by: Colangelo, Philip, et al.
Published: (2025)
by: Colangelo, Philip, et al.
Published: (2025)
Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-Chips
by: Dagli, Ismet, et al.
Published: (2023)
by: Dagli, Ismet, et al.
Published: (2023)
Equilibrium in the Computing Continuum through Active Inference
by: Sedlak, Boris, et al.
Published: (2023)
by: Sedlak, Boris, et al.
Published: (2023)
SP-IMPact: A Framework for Static Partitioning Interference Mitigation and Performance Analysis
by: Costa, Diogo, et al.
Published: (2025)
by: Costa, Diogo, et al.
Published: (2025)
Visual Insights into Agentic Optimization of Pervasive Stream Processing Services
by: Sedlak, Boris, et al.
Published: (2026)
by: Sedlak, Boris, et al.
Published: (2026)
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
by: Ng, Nathan, et al.
Published: (2026)
by: Ng, Nathan, et al.
Published: (2026)
Multi-DNN Inference of Sparse Models on Edge SoCs
by: Luo, Jiawei, et al.
Published: (2026)
by: Luo, Jiawei, et al.
Published: (2026)
Rethinking Inference Placement for Deep Learning across Edge and Cloud Platforms: A Multi-Objective Optimization Perspective and Future Directions
by: Zhang, Zongshun, et al.
Published: (2025)
by: Zhang, Zongshun, et al.
Published: (2025)
ARCAS: Adaptive Runtime System for Chiplet-Aware Scheduling
by: Fogli, Alessandro, et al.
Published: (2025)
by: Fogli, Alessandro, et al.
Published: (2025)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
by: Jeong, Bodon, et al.
Published: (2026)
by: Jeong, Bodon, et al.
Published: (2026)
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
by: Suo, Jiashun, et al.
Published: (2025)
by: Suo, Jiashun, et al.
Published: (2025)
An Uncertainty-Aware Resilience Micro-Agent for Causal Observability in the Computing Continuum
by: De Silva, Suvi, et al.
Published: (2026)
by: De Silva, Suvi, et al.
Published: (2026)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
by: Vellaisamy, Prabhu, et al.
Published: (2025)
by: Vellaisamy, Prabhu, et al.
Published: (2025)
Mitigating GIL Bottlenecks in Edge AI Systems
by: Mandal, Mridankan, et al.
Published: (2026)
by: Mandal, Mridankan, et al.
Published: (2026)
Characterizing Adaptive Mesh Refinement on Heterogeneous Platforms with Parthenon-VIBE
by: Poptani, Akash, et al.
Published: (2025)
by: Poptani, Akash, et al.
Published: (2025)
ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
by: Shen, Zixu, et al.
Published: (2025)
by: Shen, Zixu, et al.
Published: (2025)
Communication-Efficient Training Workload Balancing for Decentralized Multi-Agent Learning
by: Mohammadabadi, Seyed Mahmoud Sajjadi, et al.
Published: (2024)
by: Mohammadabadi, Seyed Mahmoud Sajjadi, et al.
Published: (2024)
Demystifying Serverless Costs on Public Platforms: Bridging Billing, Architecture, and OS Scheduling
by: Lin, Changyuan, et al.
Published: (2025)
by: Lin, Changyuan, et al.
Published: (2025)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
by: Xu, Guanyu, et al.
Published: (2025)
by: Xu, Guanyu, et al.
Published: (2025)
Evaluating Asynchronous Semantics in Trace-Discovered Resilience Models: A Case Study on the OpenTelemetry Demo
by: Krasnovsky, Anatoly A.
Published: (2025)
by: Krasnovsky, Anatoly A.
Published: (2025)
Emergence-as-Code for Self-Governing Reliable Systems
by: Krasnovsky, Anatoly A.
Published: (2026)
by: Krasnovsky, Anatoly A.
Published: (2026)
Design and Implementation of an IoT Cluster with Raspberry Pi Powered by Solar Energy: A Theoretical Approach
by: Portillo, Noel
Published: (2025)
by: Portillo, Noel
Published: (2025)
Adaptive DNN Partitioning and Offloading in Heterogeneous Edge-Cloud Continuum
by: Deng, Akuen Akoi, et al.
Published: (2026)
by: Deng, Akuen Akoi, et al.
Published: (2026)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
by: Jain, Rutwik, et al.
Published: (2026)
by: Jain, Rutwik, et al.
Published: (2026)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
by: Chakraborty, Abhinaba, et al.
Published: (2025)
by: Chakraborty, Abhinaba, et al.
Published: (2025)
Ecomap: Sustainability-Driven Optimization of Multi-Tenant DNN Execution on Edge Servers
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
EdgeProfiler: A Fast Profiling Framework for Lightweight LLMs on Edge Using Analytical Model
by: Pinnock, Alyssa, et al.
Published: (2025)
by: Pinnock, Alyssa, et al.
Published: (2025)
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
by: Scheinert, Dominik, et al.
Published: (2023)
by: Scheinert, Dominik, et al.
Published: (2023)
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
by: Wang, Yuxin, et al.
Published: (2024)
by: Wang, Yuxin, et al.
Published: (2024)
Active Inference-Based Adaptive Routing for Heterogeneous Edge AI Services
by: Wang, Zihang, et al.
Published: (2026)
by: Wang, Zihang, et al.
Published: (2026)
Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
by: Zhang, Yaozheng, et al.
Published: (2025)
by: Zhang, Yaozheng, et al.
Published: (2025)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
by: Jain, Kunal, et al.
Published: (2024)
by: Jain, Kunal, et al.
Published: (2024)
Similar Items
-
HiDP: Hierarchical DNN Partitioning for Distributed Inference on Heterogeneous Edge Platforms
by: Taufique, Zain, et al.
Published: (2024) -
Twill: Scheduling Compound AI Systems on Heterogeneous Mobile Edge Platforms
by: Taufique, Zain, et al.
Published: (2025) -
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
by: Li, Xiangchen, et al.
Published: (2025) -
Energy-Optimized Scheduling for AIoT Workloads Using TOPSIS
by: Pradeep, Preethika, et al.
Published: (2025) -
Workload composition smooths aggregate power demand while sustaining short-horizon ramps in AI data centers
by: Majumder, Subir, et al.
Published: (2026)