To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of Accelerators
Fuente:
arXiv
Saved in:
| Main Authors: | Ng, Nathan, Irwin, David, Swami, Ananthram, Towsley, Don, Shenoy, Prashant |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
by: Ng, Nathan, et al.
Published: (2026)
by: Ng, Nathan, et al.
Published: (2026)
FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
by: Wu, Li, et al.
Published: (2025)
by: Wu, Li, et al.
Published: (2025)
Proposal of Automatic Offloading Method in Mixed Offloading Destination Environment
by: Yamato, Yoji
Published: (2020)
by: Yamato, Yoji
Published: (2020)
Egret: Reinforcement Mechanism for Sequential Computation Offloading in Edge Computing
by: Peng, Haosong, et al.
Published: (2024)
by: Peng, Haosong, et al.
Published: (2024)
CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters
by: Hanafy, Walid A., et al.
Published: (2025)
by: Hanafy, Walid A., et al.
Published: (2025)
CarbonEdge: Leveraging Mesoscale Spatial Carbon-Intensity Variations for Low Carbon Edge Computing
by: Wu, Li, et al.
Published: (2025)
by: Wu, Li, et al.
Published: (2025)
OffloadFS: Leveraging Disaggregated Storage for Computation Offloading
by: Moon, Sungho, et al.
Published: (2026)
by: Moon, Sungho, et al.
Published: (2026)
Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
by: Wang, Zhibin, et al.
Published: (2025)
by: Wang, Zhibin, et al.
Published: (2025)
Resource Sharing in the Edge: A Distributed Bargaining-Theoretic Approach
by: Zafari, Faheem, et al.
Published: (2020)
by: Zafari, Faheem, et al.
Published: (2020)
FRESCO: Fast and Reliable Edge Offloading with Reputation-based Hybrid Smart Contracts
by: Zilic, Josip, et al.
Published: (2024)
by: Zilic, Josip, et al.
Published: (2024)
LLM-Enhanced Deep Reinforcement Learning for Task Offloading in Collaborative Edge Computing
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
Edge Offloading in Smart Grid
by: Arcas, Gabriel Ioan, et al.
Published: (2024)
by: Arcas, Gabriel Ioan, et al.
Published: (2024)
Distributed OpenMP Offloading of OpenMC on Intel GPU MAX Accelerators
by: Fridman, Yehonatan, et al.
Published: (2024)
by: Fridman, Yehonatan, et al.
Published: (2024)
AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
by: Kumar, Abhishek Vijaya, et al.
Published: (2024)
by: Kumar, Abhishek Vijaya, et al.
Published: (2024)
Orchestrating Joint Offloading and Scheduling for Low-Latency Edge SLAM
by: Zhang, Yao, et al.
Published: (2025)
by: Zhang, Yao, et al.
Published: (2025)
Neuro-Inspired Task Offloading in Edge-IoT Networks Using Spiking Neural Networks
by: Rossi, Fabio Diniz
Published: (2025)
by: Rossi, Fabio Diniz
Published: (2025)
Untangling Carbon-free Energy Attribution and Carbon Intensity Estimation for Carbon-aware Computing
by: Maji, Diptyaroop, et al.
Published: (2023)
by: Maji, Diptyaroop, et al.
Published: (2023)
The Green Mirage: Impact of Location- and Market-based Carbon Intensity Estimation on Carbon Optimization Efficacy
by: Maji, Diptyaroop, et al.
Published: (2024)
by: Maji, Diptyaroop, et al.
Published: (2024)
Plug & Offload: Transparently Offloading TCP Stack onto Off-path SmartNIC with PnO-TCP
by: Nan, Hailong, et al.
Published: (2025)
by: Nan, Hailong, et al.
Published: (2025)
Disaggregated Memory with SmartNIC Offloading: a Case Study on Graph Processing
by: Wahlgren, Jacob, et al.
Published: (2024)
by: Wahlgren, Jacob, et al.
Published: (2024)
Collaborative Inference for Large Models with Task Offloading and Early Exiting
by: Xie, Zuan, et al.
Published: (2024)
by: Xie, Zuan, et al.
Published: (2024)
A Survey of Computation Offloading with Task Types
by: Zhang, Siqi, et al.
Published: (2023)
by: Zhang, Siqi, et al.
Published: (2023)
Energy-Efficient Joint Offloading and Resource Allocation for Deadline-Constrained Tasks in Multi-Access Edge Computing
by: Gao, Chuanchao, et al.
Published: (2025)
by: Gao, Chuanchao, et al.
Published: (2025)
Workload Distribution with Rateless Encoding: A Low-Latency Computation Offloading Method within Edge Networks
by: Guo, Zhongfu, et al.
Published: (2023)
by: Guo, Zhongfu, et al.
Published: (2023)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
by: Yang, Zheming, et al.
Published: (2026)
by: Yang, Zheming, et al.
Published: (2026)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
by: Ma, Chenxiang, et al.
Published: (2025)
by: Ma, Chenxiang, et al.
Published: (2025)
Decentralized Proactive Model Offloading and Resource Allocation for Split and Federated Learning
by: Huang, Binbin, et al.
Published: (2024)
by: Huang, Binbin, et al.
Published: (2024)
A Joint Time and Energy-Efficient Federated Learning-based Computation Offloading Method for Mobile Edge Computing
by: Mukherjee, Anwesha, et al.
Published: (2024)
by: Mukherjee, Anwesha, et al.
Published: (2024)
Blockchain-Enhanced Offloading in Mobile Edge Computing: A Systematic Review and Survey of Current Trends and Future Directions
by: Moghaddasi, Komeil, et al.
Published: (2024)
by: Moghaddasi, Komeil, et al.
Published: (2024)
A Model Aware AIGC Task Offloading Algorithm in IIoT Edge Computing
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Rethinking Inter-Process Communication with Memory Operation Offloading
by: Park, Misun, et al.
Published: (2026)
by: Park, Misun, et al.
Published: (2026)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
by: Zhang, Huawei, et al.
Published: (2025)
by: Zhang, Huawei, et al.
Published: (2025)
Preemption Aware Task Scheduling for Priority and Deadline Constrained DNN Inference Task Offloading in Homogeneous Mobile-Edge Networks
by: Cotter, Jamie, et al.
Published: (2025)
by: Cotter, Jamie, et al.
Published: (2025)
MoA-Off: Adaptive Heterogeneous Modality-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
by: Yang, Zheming, et al.
Published: (2025)
by: Yang, Zheming, et al.
Published: (2025)
Joint Optimization of Offloading, Batching and DVFS for Multiuser Co-Inference
by: Xu, Yaodan, et al.
Published: (2025)
by: Xu, Yaodan, et al.
Published: (2025)
Communication Offloading on SmartNIC DPUs: A Quantitative Approach
by: Wahlgren, Jacob, et al.
Published: (2026)
by: Wahlgren, Jacob, et al.
Published: (2026)
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
by: Khalilov, Mikhail, et al.
Published: (2024)
by: Khalilov, Mikhail, et al.
Published: (2024)
Static Generation of Efficient OpenMP Offload Data Mappings
by: Marzen, Luke, et al.
Published: (2024)
by: Marzen, Luke, et al.
Published: (2024)
TURNIP: A "Nondeterministic" GPU Runtime with CPU RAM Offload
by: Ding, Zhimin, et al.
Published: (2024)
by: Ding, Zhimin, et al.
Published: (2024)
SLO-Aware Task Offloading within Collaborative Vehicle Platoons
by: Sedlak, Boris, et al.
Published: (2024)
by: Sedlak, Boris, et al.
Published: (2024)
Similar Items
-
Collaborative Processing for Multi-Tenant Inference on Memory-Constrained Edge TPUs
by: Ng, Nathan, et al.
Published: (2026) -
FailLite: Failure-Resilient Model Serving for Resource-Constrained Edge Environments
by: Wu, Li, et al.
Published: (2025) -
Proposal of Automatic Offloading Method in Mixed Offloading Destination Environment
by: Yamato, Yoji
Published: (2020) -
Egret: Reinforcement Mechanism for Sequential Computation Offloading in Edge Computing
by: Peng, Haosong, et al.
Published: (2024) -
CarbonFlex: Enabling Carbon-aware Provisioning and Scheduling for Cloud Clusters
by: Hanafy, Walid A., et al.
Published: (2025)